Autonomous Driving Control With Closed-Loop Decision and Planning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current autonomous driving systems face challenges in further improving traveling safety due to limitations in the decision-making and control processes between the behavior decision-making layer and the motion planning layer.
Innovation Solution
A method for closed-loop optimization of the behavior decision-making layer and motion planning layer is introduced, where both layers are optimized based on differences between their outputs and a target teaching traveling sequence, using a determining model to refine their performance and learn optimal driving behaviors and trajectories.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the behavior decision-making layer and motion planning layer are optimized separately using conventional feedback methods, then the system complexity is reduced and ease of operation is improved, but the traveling safety and performance of the autonomous driving system cannot be further improved
Solution Approach 1:
The patent merges the behavior decision-making layer and motion planning layer into a unified optimization framework. Both layers are optimized simultaneously using a joint loss function that incorporates objectives from both layers, allowing them to work together as an integrated system rather than separate modules. This combining approach enables improved safety performance while managing system complexity through unified training.
Solution Approach 2:
The patent introduces a reward model as an intermediary component that evaluates and guides the optimization of both the behavior decision-making layer and motion planning layer. This reward model acts as a mediator that provides feedback signals to both layers, enabling coordinated optimization without directly increasing the complexity of either individual layer. The reward model bridges the two layers and facilitates their joint improvement in safety performance.
2Reliability
If the autonomous driving system uses a hierarchical architecture with behavior decision-making layer and motion planning layer, then the adaptability and decision-making capability are improved, but the system cannot achieve optimal performance in terms of traveling safety
Solution Approach 1:
The patent implements a feedback mechanism where a reward model evaluates the outputs of both the behavior decision-making layer and motion planning layer. The reward model provides feedback signals that guide the optimization of both layers, allowing the system to learn from evaluation results and continuously improve safety performance. This feedback loop enables the hierarchical architecture to achieve optimal performance by iteratively refining decisions based on evaluated outcomes.
Solution Approach 2:
The patent makes the optimization process dynamic by using a reward model that can adaptively evaluate different scenarios and provide context-dependent feedback. The reward model allows the system to dynamically adjust its optimization priorities based on the specific driving situation, enabling the hierarchical architecture to achieve better safety performance across diverse scenarios rather than relying on static optimization parameters.
Data Source
AI summary
In the method for optimizing decision-making regulation and control, a first traveling sequence is obtained, where the first traveling sequence includes a first trajectory sequence of the vehicle in information about a first environment and first target driving behavior output by a behavior decision-making layer of a decision-making and control system based on the information about the first environment. A second traveling sequence is obtained, where the second traveling sequence includes a second trajectory sequence output by a motion planning layer of the decision-making and control system based on preset second target driving behavior and the second target driving behavior. The behavior decision-making layer is optimized based on a difference between the first traveling sequence and a preset traveling sequence, and the motion planning layer is optimized based on a difference between the second traveling sequence and the preset traveling sequence.


