Adaptive Action Planning Hyperparameters for Automated Driving
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current action planning systems for automated driving face challenges in handling diverse environment and vehicle conditions, leading to computational inefficiencies and sub-optimal decision-making in scenarios like congestion, inclement weather, and unmodeled road conditions.
Innovation Solution
A system and method that utilizes a reinforcement learning agent with a planning policy to adaptively tune hyperparameters based on sensor and lane data, adjusting hyperparameters using activation functions to optimize trajectory actions and enhance computational efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If current online action planning systems use fixed engineered hyperparameters, then the system structure is simple, but computational efficiency deteriorates and time efficiency is compromised in disparate scenarios
Solution Approach 1:
The patent applies dynamics by transforming fixed hyperparameters into dynamic, adaptive parameters. The system uses reinforcement learning agents that continuously adjust hyperparameters based on real-time environmental conditions, traffic scenarios, and vehicle states. This allows the planning system to adapt its computational behavior dynamically, improving efficiency in disparate scenarios without requiring a completely complex system architecture.
Solution Approach 2:
The patent directly implements parameter changes by modifying hyperparameters such as planning horizon, resolution, and search depth based on detected driving conditions. The reinforcement learning agent learns optimal parameter configurations for different scenarios (e.g., congested traffic vs. open roads), enabling the system to balance computational efficiency with performance by changing parameters adaptively rather than using fixed values.
2Adaptability or versatility
If action planning systems use fixed hyperparameters, then the system is easy to operate, but adaptability to diverse environment conditions deteriorates
Solution Approach 1:
The patent implements self-service by enabling the system to automatically adjust its own hyperparameters without external intervention. The reinforcement learning agents monitor environmental conditions and autonomously modify planning parameters to optimize performance for the current driving scenario, making the system adaptive while maintaining ease of operation through automation.
Solution Approach 2:
The system uses feedback mechanisms where the reinforcement learning agents continuously monitor planning outcomes and environmental conditions, then adjust hyperparameters based on this feedback. This closed-loop approach enables the system to adapt to diverse driving conditions by learning from past performance and modifying parameters accordingly, while the automation maintains operational simplicity.
3Loss of time
If online action planning is performed with fixed parameters, then computational resources are conserved, but time efficiency is compromised in congested situations and unmodeled road conditions
Solution Approach 1:
The patent applies parameter changes by dynamically adjusting planning parameters such as time horizon, spatial resolution, and search depth based on detected driving conditions. In congested or complex scenarios, the system increases computational resources allocated to planning; in simpler scenarios, it reduces them. This adaptive parameter adjustment improves time efficiency when needed while managing computational resource consumption overall.
Data Source
AI summary
A method of adaptively tuning parameters in action planning for automated driving of a vehicle to a destination is provided. The method comprises receiving sensor data from a vehicle sensor, lane data of a road plan to the destination, and a plurality of first hyperparameters in a reinforcement learning agent at an initial state and an initial timestamp. The reinforcement learning agent has a planning policy. The method comprises adjusting the plurality of first hyperparameters via the planning policy having at least one first activation function to define an output defining a plurality of second hyperparameters. The method comprises determining a baseline trajectory action based on a trajectory reward value at a final state and a final timestamp. The method comprises modifying the baseline trajectory action between the initial state and the final state to define a refined trajectory action and controlling the vehicle based on the refined trajectory action.

