Autonomous Driving Solution in Multi-Objective Complex Traffic Scenarios Based on Reinforcement Learning

Through the combination of reinforcement learning and rule constraints, the universality and safety of autonomous driving in multi-objective complex traffic scenarios are solved, and efficient and safe autonomous driving in multiple traffic scenarios is achieved.

CN114701517BActive Publication Date: 2025-07-22NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210370991.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-07-22
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

The existing autonomous driving technology has poor generalization and insufficient safety when dealing with multi-objective complex traffic scenarios, and there are generalization and stability problems in deep reinforcement learning, making it difficult to achieve completely unmanned driving.

Method used

Using reinforcement learning-based methods, through anthropomorphic observation design and reward reshaping, combined with the hazardous action recognizer of the LSTM network and the knowledge-based rule constraint, an autonomous driving model in multi-objective complex traffic scenarios is constructed, and the model's performance in different scenarios is improved using time-varying training strategies.

Benefits of technology

It has achieved good universality and generalization performance in various traffic scenarios, improved the safety and stability of autonomous driving, reduced the number of collisions, and ensured the safety of vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114701517B_ABST
    Figure CN114701517B_ABST
Patent Text Reader

Abstract

The present invention discloses a solution for autonomous driving in multi-objective complex traffic scenarios based on reinforcement learning. This method can use a set of reinforcement learning autonomous driving modeling methods to handle all traffic scenarios, with good generality. The comprehensive reinforcement learning modeling is based on the traditional reinforcement learning framework, using environmental perception information and feature quantities extracted by combining human knowledge as the observation space. The model training is based on a time-varying training strategy to improve the training speed and the generalization of policy application. To further ensure its formal safety, a dangerous action recognizer based on the long short-term memory (LSTM) network and a rule constraint based on the human knowledge system are also proposed. The dangerous action recognizer is sampled and trained from the environment, enabling the vehicle to have the ability to recognize dangerous actions and dangerous scenarios. And rule constraints are designed for specific situations to limit the output actions, which can greatly improve safety and reduce the number of collisions to ensure the driving safety of the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for autonomous driving in a multi-objective complex traffic scenario based on reinforcement learning, belonging to the technical field of autonomous driving, and particularly relates to a general autonomous driving algorithm modeling and training scheme based on deep reinforcement learning in a multi-objective complex traffic scenario. Background Art

[0002] With the rapid development of the intelligent vehicle industry and the continuous maturity of autonomous driving technology, driverless technology has become the future trend of vehicle development. Currently, the implemented autonomous driving systems can reach the L3 level, that is, the vehicle can automatically drive under the condition of driver monitoring, and the driver needs to take over the vehicle in case of an emergency. However, one of the important reasons why fully driverless has not been achieved is that the current rule-based decision-making methods cannot handle sufficient traffic scenarios, and there are significant potential safety risks.

[0003] The mainstream autonomous driving technologies can be divided into three modules: perception, decision-making, and control. Among them, the decision-making module is the core part of the intelligent system. Currently, autonomous driving decision-making technologies can be mainly divided into two categories: rule-based and learning-based. Rule-based decision-making technologies include general decision-making models, finite state machine models, decision tree models, knowledge reasoning-based models, etc. Learning-based decision-making technologies are mainly based on deep learning and reinforcement learning.

[0004] Currently, rule-based decision-making technologies are still in actual use, but more and more problems have emerged. Rule-based systems are difficult to enumerate all possible scenarios, and traffic accidents are likely to occur in some unconsidered scenarios. Secondly, the design of rule systems is relatively costly in terms of human resources and system complexity, and system maintenance and upgrade are particularly cumbersome. Therefore, there is an urgent need to develop and improve other technical methods, and data-driven deep reinforcement learning is one direction. However, the current application of deep reinforcement learning in autonomous driving is mainly in a specific scenario, such as overtaking, lane changing, lane keeping, etc., and its generality is not strong. In addition, deep neural networks currently do not have complete interpretability, have the problem of catastrophic forgetting, and are prone to generating some unknown unsafe actions. On the other hand, reinforcement learning itself also has problems of generalization and stability. Therefore, to make deep reinforcement learning have practical applications in autonomous driving decision control can, to a certain extent, improve and solve the problems of poor generality, weak generalization ability, and unimproved safety. Summary of the Invention

[0005] The content of this application is partially used to introduce concepts in a brief form, which will be described in detail in the following detailed implementation section. The content section of this application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Aiming at the problems and deficiencies in the prior art, the purpose of the present invention is to provide a solution for autonomous driving in a multi-objective complex traffic scenario based on reinforcement learning. Through anthropomorphic observation design and reward function design based on reward reshaping, it can be achieved that only one model can exhibit better comprehensive performance in traffic scenarios with multiple types of environments having multiple strategies. In addition, in order to improve the training speed and generalization, the present invention proposes a reward time-varying training method. In order to enhance safety, the present invention proposes methods for ensuring safety such as an LSTM-based dangerous action recognizer and knowledge-based safety filtering to solve the problems raised in the above background technology.

[0007] To achieve the above purpose, the present invention provides the following technical solutions:

[0008] The present invention discloses a solution for autonomous driving in a multi-objective complex traffic scenario based on reinforcement learning, including the following steps:

[0009] Step 1, prepare a simulator environment for autonomous driving simulation and a complex driving scenario;

[0010] Step 2, add environmental feature information required for training the reinforcement learning model as environmental observation information in the observation space, including ego vehicle information, other vehicle information, and road information, and calculate key feature information according to the environmental feature information, including the time to collision of each lane with the vehicle in front, the time difference of the ego vehicle's distance to the collision point, and the variance of the orientation angle between the ego vehicle and the front road point;

[0011] Step 3, set the reward framework required for training the reinforcement learning model;

[0012] Step 4, train the reinforcement learning model based on the time-varying training method, and then continue to train in different traffic scenarios by combining with the use of a meta-model. Modify the reward weights during training according to the driving performance of the agent and the types of collisions that occur every certain number of iteration rounds, and repeat modifying the weights multiple times to end the training;

[0013] Step 5, after outputting the trained reinforcement learning model, construct a dangerous action recognizer and a rule constraint device, judge the degree of scene danger according to the environmental observation information, limit or adjust the planned quantity output by the reinforcement learning model, and continuously manually add and optimize rules by observing the effects in the simulation environment.

[0014] Further, environmental feature information required for training the reinforcement learning model is added to the observation space in step 2. The ego-vehicle information includes the ego-vehicle speed, the steering wheel angle of the ego-vehicle, the lane number where the ego-vehicle is located, the distance between the center point of the ego-vehicle and the center line of the lane where it is located, the deviation between the fifteen road points ahead starting from the ego-vehicle position and the heading angle of the ego-vehicle, and the relative position to the preview point at a distance of 4 / 8 meters. The other-vehicle information includes the distance of the neighboring vehicle in each lane, the time to collision with the vehicle ahead in each lane, and the relative speed between the nearest neighboring other-vehicle and the ego-vehicle in the ego-vehicle lane. The road information includes the lateral distance of the ego-vehicle from the center of the lane and the heading error of the road point.

[0015] Further, the reward framework in step 3 includes an environmental reward, a speed reward, a collision penalty, and a lane center deviation penalty. Among them, the environmental reward is the survival time of the ego-vehicle, representing the duration from the starting point to the collision of the ego-vehicle. The speed reward is the driving speed of the ego-vehicle, measured in meters traveled per second. The collision penalty is given when the ego-vehicle leaves the route, drives out of the road boundary, or collides with an environmental vehicle. The lane center deviation penalty is the absolute value of the distance between the vehicle center and the center line.

[0016] Further, in step 4, the reinforcement learning model is trained by combining the time-varying training method with the meta-model. The specific steps are as follows:

[0017] Step 4.1, initialize the reinforcement learning model, and use each scenario to train for a certain number of rounds in sequence to obtain a meta-model.

[0018] Step 4.2, use the time-varying training method to train the meta-model obtained in step 4.1 in the selected scenario, and adjust the reward weights according to the defects existing in the agent's behavior.

[0019] Step 4.3, set the scenario to all simple scenarios without intersections, and repeat the process of step 4.2 to train and improve the performance in simple scenarios without intersections.

[0020] Step 4.4, set the scenario to all scenarios with intersections, and repeat the process of step 4.2 to train and improve the performance in scenarios with intersections.

[0021] Step 4.5, set the scenario to scenarios with roundabouts and multi-directional vehicles, and repeat the process of step 4.2 to train and improve the performance in scenarios with roundabouts and multi-directional vehicles.

[0022] Step 4.6, continue training in the remaining scenarios until the process ends.

[0023] Further, the specific steps for training the reinforcement learning model by the time-varying training method in step 4 are as follows:

[0024] Step 4.2.1, set the hyperparameters of the reinforcement learning model;

[0025] Step 4.2.2, set the reward function as the basic reward, enable the Agent to learn lane keeping, and start iterative training;

[0026] Step 4.2.3, increase the weights of the lane center deviation penalty and the collision penalty, and continue iterative training;

[0027] Step 4.2.4, continue to increase the collision penalty and continue iterative training;

[0028] Step 4.2.5, add new scenarios on the basis of the original scenario dataset, add speed rewards, and increase the weights of the lane center deviation penalty and the collision penalty until the iteration ends.

[0029] Further, in step 5, the dangerous action recognizer predicts its danger level according to the actions and environmental observation information output by the reinforcement learning model, and takes actions such as emergency avoidance and emergency adjustment according to the danger level. The dangerous action recognizer includes a sample collection stage and a training stage. The specific steps of the sample collection stage are as follows:

[0030] Step 5.1.1, prepare various types of scenarios and select a scenario to start training;

[0031] Step 5.1.2, initialize the PPO policy model and start training the policy model in the selected scenario;

[0032] Step 5.1.3, record the trajectory during this round of operation during the running process;

[0033] Step 5.1.4, when a collision occurs, collect the 10 steps before the collision as negative samples, and randomly collect any continuous 10 steps in this round of trajectory as positive samples;

[0034] Step 5.1.5, after training until the running steps reach the set total number of steps, select the next scenario to repeat training;

[0035] Step 5.1.6, end until all scenarios are collected.

[0036] Further, in step 5, the dangerous action recognizer is a dangerous action recognizer model based on a long short-term memory network. The specific steps of the training stage are as follows:

[0037] Step 5.2.1, according to each group of collected sample data, use a sliding window to generate data groups as the model input, and use the label of the last step of each group as the target label of the model;

[0038] Step 5.2.2: Use the Adam optimizer and adjust the optimizer learning rate based on the cosine simulated annealing method;

[0039] Step 5.2.3: Use the mean squared error loss as the training loss function and calculate the mean squared error loss between the model output and the target label;

[0040] Step 5.2.4: Set relevant model parameters and the number of training model rounds to complete the training.

[0041] Furthermore, in Step 5, the rule constraint is written based on human knowledge and empirical statistics of simulation experiments to restrict the behavior of the ego vehicle in certain specific situations. The rule constraint mainly includes knowledge rules for closest distance protection, knowledge rules at intersections, knowledge rules before sharp turns, and correction rules for long-term staying when there are no neighboring vehicles around the ego vehicle. The environmental observation information judges different scenarios respectively to determine the speed limit rule of the ego vehicle.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a solution for autonomous driving in multi-objective complex traffic scenarios based on reinforcement learning. This method can use a set of reinforcement learning autonomous driving modeling methods to handle all traffic scenarios, has good generality, and can obtain good multi-objective performance and generalization performance. The reinforcement learning comprehensive modeling is based on the traditional reinforcement learning framework, uses environmental perception information and feature quantities extracted by combining human knowledge as the observation space, and sets lane keeping, driving distance, collision avoidance, etc. as rewards and punishments for the intelligent vehicle in the reinforcement learning algorithm according to the evaluation index. When training the model, the meta-learning idea is combined with the time-varying training strategy. Different reward weights and different training sets are set in each stage to strengthen the partial behavioral defects formed by the intelligent agent in the previous training stage and improve the performance of the intelligent agent in some weak scenarios, which can improve the training speed and the generalization of policy application. In addition, to further ensure its formal safety, a dangerous action recognizer based on the long short-term memory (LSTM) network and a rule constraint based on the human knowledge system are also proposed. The dangerous action recognizer is sampled from the environment and trained to enable the vehicle to have the ability to recognize dangerous actions and dangerous scenarios, and rule constraints are designed for specific situations to limit the output actions, which can greatly improve safety, reduce the number of collisions, and handle special emergencies to ensure the driving safety of the vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings that form a part of this application are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and descriptions of the accompanying drawings of this application are used to explain this application and do not constitute an improper limitation of this application.

[0044] In the drawings:

[0045] Figure 1 This is the overall flowchart of the autonomous driving solution method in a multi-objective complex traffic scenario based on reinforcement learning of the present invention;

[0046] Figure 2 This is the step diagram of the autonomous driving solution method in a multi-objective complex traffic scenario based on reinforcement learning of the present invention;

[0047] Figure 3 This is the flowchart of the time-varying training method and the meta-model training reinforcement learning model in the reinforcement learning of the present invention;

[0048] Figure 4 This is the framework diagram of combining end-to-end and rule constraint to control autonomous driving of the present invention;

[0049] Figure 5 This is the schematic diagram of the reinforcement learning policy network based on proximal policy optimization of the present invention;

[0050] Figure 6 This is the schematic diagram of the model structure of the dangerous action recognizer based on the long short-term memory (LSTM) neural network of the present invention;

[0051] Figure 7 This is the flowchart of the sampling stage of the dangerous action recognizer of the present invention;

[0052] Figure 8 This is the computer language block diagram of the sampling algorithm of the dangerous action recognizer of the present invention. Detailed implementation manners

[0053] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0054] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0055] The present invention discloses an autonomous driving solution method in a multi-objective complex traffic scenario based on reinforcement learning. Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings and in combination with embodiments. Referring to Figures 1 to 2 as shown, it mainly includes the following steps:

[0056] Step 1: Prepare a simulator environment for autonomous driving simulation and a complex driving scenario;

[0057] Step 2, add the environmental feature information required for training the reinforcement learning model as environmental observation information in the observation space, including ego-vehicle information, other-vehicle information, and road information, and calculate the key feature information according to the environmental feature information, including the time to collision of each lane with the vehicle in front, the time difference of the ego-vehicle from the collision point, and the variance of the orientation angle between the ego-vehicle and the front road point;

[0058] Step 3, set the reward framework required for training the reinforcement learning model;

[0059] Step 4, train the reinforcement learning model based on the time-varying training method, and then continue to train in different traffic scenarios by using the meta-model. Modify the reward weights during training every certain number of iteration rounds according to the driving performance of the agent and the types of collisions that occur, and repeat modifying the weights multiple times to end the training;

[0060] Step 5, after outputting the trained reinforcement learning model, construct a dangerous action recognizer and a rule constraint device, judge the degree of scene danger according to the environmental observation information, limit or adjust the planned quantity output by the reinforcement learning model, and continuously manually add and optimize rules by observing the effects in the simulation environment.

[0061] The multi-objectives of the present invention refer to the goals of precise vehicle driving, high driving speed, driving safety, high algorithm robustness, and high generalization. Specifically, in a single traffic scenario, when the same number of simulations are performed under the condition of the maximum driving distance of the vehicle, the driving speed is reflected in that the average speed of the ego-vehicle is as large as possible. Driving safety is reflected in that the number of collisions between the ego-vehicle and the environmental vehicles, the number of times the ego-vehicle deviates from the route, and the driving distance of the ego-vehicle are as small as possible. Driving precision is reflected in that the average distance between the center point of the ego-vehicle and the road center line is as close as possible. High algorithm robustness and high generalization are reflected in that an algorithm can also achieve good results in multiple complex traffic scenarios and map scenarios that have not been trained.

[0062] The complex traffic scenario means that the map scenarios used in the simulator training have various road types and traffic flow conditions. The map types used include nine map types such as simple, sharp bend, intersection, roundabout, merging, diverging, and mixed road scenarios. Different scenarios also include environmental vehicles with different densities. In the rules set in the simulation environment, the driving trajectories and driving strategies of environmental vehicles have a certain degree of randomness. The strategy types can be divided into three types: conservative, moderate, and aggressive. Under each strategy, the operation strategy parameters of environmental vehicles still have a certain degree of randomness.

[0063] In step 2, environmental feature information required for training the reinforcement learning model is added to the observation space as environmental observation information, and key feature quantities are calculated based on the environmental feature information. In this reinforcement learning model, the input observation space includes environmental perception information such as ego-vehicle information, other-vehicle information, and road information, as well as observation features extracted from the environmental information; its output actions include throttle opening, brake control, and steering wheel angle control.

[0064] Specifically, the environmental observation information includes environmental perception information such as ego-vehicle information, other-vehicle information, and road information. Among them, the ego-vehicle information includes ego-vehicle speed, ego-vehicle steering wheel angle, the sequence number of the lane where the ego-vehicle is located, the distance between the center point of the ego-vehicle and the center line of the lane, the deviation of the first fifteen road points in front of the ego-vehicle from the ego-vehicle's heading angle starting from the ego-vehicle's position, and the relative position to the preview point at a distance of 4 / 8 meters, totaling 2 + 1 + 1 + 15 + 4 = 19 dimensions. The other-vehicle information includes the distance of the neighboring vehicle in each lane, the time to collision (TTC) with the vehicle in front in each lane, and the relative speed between the nearest neighboring other-vehicle in the ego-vehicle's lane and the ego-vehicle, totaling 6 * 3 = 18 dimensions; when the ego-vehicle travels to an intersection, the relative position, relative heading, relative speed, and time difference of collision of the five vehicles closest to the ego-vehicle and facing the intersection within a range of fifty meters are queried, totaling 5 * 5 = 25 dimensions. The road information includes: the lateral distance of the ego-vehicle from the center of the lane, and the heading error of the road point. The heading error of the road point represents the distance between the straight line where the ego-vehicle is heading and the road point in front. Starting from the road point at the ego-vehicle's position, the heading errors of the first fifteen road points in front are calculated respectively. Based on the principle that the road points closer to the ego-vehicle are denser and the road points farther away are sparser, the number of selected road points is [0, 1, 2, 3, 5, 7, 10, 13, 17, 21, 25, 30, 35, 42, 50], representing the i-th road point in front and also the distance of the road point from the intelligent vehicle.

[0065] The extracted key feature information is mainly collision avoidance vehicle information. The three other-vehicles closest to the ego-vehicle's position are selected as collision avoidance vehicles, and the other-vehicle information is calculated in the ego-vehicle coordinate system. Therefore, the key feature information includes the relative position of the collision avoidance vehicle, the absolute speed of the collision avoidance vehicle, the relative heading angle between the ego-vehicle and the collision avoidance vehicle, and the time difference of collision (TDTC) from the collision point. The calculation formula for the time difference of collision (TDTC) from the collision point is:

[0066]

[0067] In step 3, set up the reward framework required for training the reinforcement learning model. In the reward framework for training the reinforcement learning model, it includes environment reward, speed reward, collision penalty, and lane center deviation penalty. Among them, the environment reward is the survival time of the ego vehicle, which represents the duration from the starting point of the ego vehicle to the occurrence of a collision. Its value gradually increases from 1 to 4, and then starts to increase step by step from 1. In the simulation environment, as long as the ego vehicle is still alive in a single-step simulation, a reward is given. The speed reward is the driving speed of the ego vehicle, with the unit of the distance traveled per second. The lane center deviation penalty is the absolute value of the distance between the vehicle center and the center line. The collision penalty includes three cases. When the ego vehicle drives out of the route, drives out of the road boundary, or collides with an environmental vehicle, a corresponding collision penalty will be given. Its value is a constant 5, and its weight increases with the number of iterations. That is, a penalty of -5 is obtained for each case, and the coefficients of each penalty are 1, 0.5, 0.1, and 1.2 respectively. These coefficients will be adjusted during the training process.

[0068] Refer to Figure 3 As shown, based on the time-varying training method combined with the meta-model to train the reinforcement learning model, the performance of the vehicle in some special scenarios can be improved, such as intersections, roundabouts, etc. The following gives its training steps:

[0069] Step 4.1, initialize the reinforcement learning model, and use each scenario to train for a certain number of rounds in turn to obtain a meta-model;

[0070] Step 4.2, use the time-varying training method to train the meta-model obtained in step 4.1 in the selected scenario, and adjust the reward weights according to the defects existing in the agent's behavior;

[0071] Step 4.3, set the scenario as all simple scenarios without intersections, and repeat the process of step 4.2 for training to improve the performance in simple scenarios without intersections;

[0072] Step 4.4, set the scenario as all scenarios with intersections, and repeat the process of step 4.2 for training to improve the performance in intersection scenarios;

[0073] Step 4.5, set the scenario as the scenario with roundabouts and multi-directional vehicles, and repeat the process of step 4.2 for training to improve the performance in scenarios with roundabouts and multi-directional vehicles;

[0074] Step 4.6, continue training in the remaining scenarios until the process ends.

[0075] Specifically, in step 4.2, the specific process of using the time-varying training method to train the reinforcement learning model is as follows:

[0076] Step 4.2.1, set the hyperparameters of the reinforcement learning model;

[0077] Step 4.2.2, set the reward function to the basic reward, enable the Agent to learn lane keeping, and start iterative training;

[0078] Step 4.2.3, increase the weights of the lane center deviation penalty and the collision penalty, and continue iterative training;

[0079] Step 4.2.4, continue to increase the collision penalty and continue iterative training;

[0080] Step 4.2.5, add new scenarios on the basis of the original scenario dataset, add speed rewards, and increase the weights of the lane center deviation penalty and the collision penalty until the iterative training ends.

[0081] The hyperparameters of the reinforcement learning model set are shown in the following table:

[0082] Parameter Name Parameter Value Learning Rate 1e-4 Training Batch Size 10240*3 λ 0.95 Clip Parameter 0.2 Number of SGD Iterations 10 SGD Mini-batch Size 1024 Step Range 1000

[0083] First, set the reward function to the basic function, enable the Agent to learn lane keeping, and express the reward function as: 1.0 × environmental reward + 0.1 × lane center deviation penalty + 1.2 × collision penalty, and start iterative training for 460 rounds from the 0th round. Then, increase the weights of the lane center deviation penalty and the collision penalty. At this time, the reward function is expressed as: 1.0 × environmental reward + 0.4 × lane center deviation penalty + 1.6 × collision penalty, and continue training from the 461st round to the 768th round. Continue to increase the collision penalty weight. At this time, the reward function is expressed as: 1.0 × environmental reward + 0.4 × lane center deviation penalty + 1.8 × collision penalty, and continue training from the 769th round to the 1152nd round. Finally, add 3 new scenarios on the basis of the original scenario dataset, including 2 all_loop scenarios and 1 mix_loop scenario, add speed rewards (the faster the vehicle travels, the greater the reward value), and increase the weights of the lane center deviation penalty and the collision penalty. At this time, it is expressed as: 1.0 × environmental reward + 0.4 × speed reward + 0.56 × lane center deviation penalty + 2.9 × collision penalty, and continue training from the 1153rd round to the 1400th round, and the training process ends.

[0084] Based on traditional reinforcement learning, the reinforcement learning model adopts time-varying strategy training to generate an agent model with higher robustness and generalization ability. The time-varying strategy of the reinforcement learning model is a phased learning method. The Agent task objectives are divided into different subtasks according to their importance and learned in different stages. In each stage, the learning tasks are hierarchically reinforced by adjusting the reward weights. At the beginning, the Agent is trained to have the ability to keep in the lane, then its collision avoidance ability, and finally its high-speed driving ability. The training process of the intelligent vehicle model consists of multiple iterations. In each iteration, the reward weights for this iteration are adjusted according to the previous simulation results. Specifically, in each iteration, the collisions and their types that occurred in the previous simulation are recorded, and the proportion of the reward is adjusted according to the proportion of a certain type of collision in all collisions.

[0085] In step 5, the dangerous action recognizer includes two stages: sample collection and training. Refer to Figure 7 and Figure 8 As shown, the specific process of sample collection is as follows:

[0086] Step 5.1.1, prepare various types of scenarios and select a scenario to start training;

[0087] Step 5.1.2, initialize the PPO policy model and start training the policy model in the selected scenario;

[0088] Step 5.1.3, record the trajectory during this round of operation during the running process;

[0089] Step 5.1.4, when a collision occurs, collect the 10 steps before the collision as negative samples, and randomly collect any consecutive 10 steps in this round of trajectory as positive samples;

[0090] Step 5.1.5, after training until the running steps reach the set total number of steps, select the next scenario to repeat the training;

[0091] Step 5.1.6, end until all scenarios are collected.

[0092] Specifically, here we use 7 categories of scenarios and select one scenario to start training. Initialize the PPO policy model, train the policy from scratch under the selected scenario, and set the total number of running steps to 400,000 steps. During the running process, record the trajectory of this round of running, including the observation-action pairs (s1, a1, s2, a2, …) at each step and the running reward values (r1, r2, …). If a collision occurs, record the observation s-action a-reward r pairs in the 10 steps before the collision as negative samples for collection, and record the trajectory length of this round as m. Randomly select a random number k from the interval [1, m - 19], and collect the observation s-action a-reward r pairs in 10 steps starting from the kth step as positive samples for collection. Calculate the 10 rolling averages using the reward values in each group of collected samples as the training labels for this group of samples. The calculation formula is as follows:

[0093]

[0094] where i represents the i-th scenario and j represents the j-th step in the 10-step sample. Store the observation s-action a-label pairs as the training dataset and train until the running steps reach the set total number of 400,000 steps. Select the next scenario and repeat steps 5.1.2 to 5.1.5 for training until all scenarios are collected and ended. The total number of collected samples is approximately 100,000.

[0095] Refer to Figure 5 and Figure 6 As shown, build a dangerous action recognizer model based on the long short-term memory (LSTM) network. The model includes an LSTM layer and 2 fully connected layers. After the output of the long short-term memory (LSTM) network passes through the first fully connected layer, the ReLU activation function is used and then connected to the second fully connected layer. Then the training process of the dangerous action recognizer is as follows:

[0096] Step 5.2.1, according to each group of sample data collected, use the sliding window to generate data groups as the model input, and use the label of the last step in each group as the target label of the model;

[0097] Step 5.2.2, use the Adam optimizer and adjust the learning rate of the optimizer based on the cosine annealing method;

[0098] Step 5.2.3, use the mean square error loss as the training loss function and calculate the mean square error loss between the model output and the target label;

[0099] Step 5.2.4, set the relevant model parameters and the number of training model rounds to complete the training.

[0100] Specifically, each group of collected samples contains 10-step observation s-action a-label pairs. Among them, the observation is 44-dimensional, the action is 3-dimensional. After concatenating the observation and the action into a 47-dimensional vector, the 10-step data is used to generate 6 groups of data with a length of 5 using a sliding window of size 5. Each group of 5-step data is used as the model input, and the label of the last step of each group is used as the target label of the model. The Adam optimizer is used, and the learning rate of the optimizer is adjusted based on the cosine annealing method. The mean squared error loss is used as the training loss function to calculate the mean squared error loss between the model output and the target label. Set the model parameters as shown in the following table and end the training after 100 rounds of the model.

[0101] Parameter Name Parameter Value LSTM Hidden Layer Size 128 Fully Connected Layer 1 128x64 Fully Connected Layer 2 64x1 Initial Learning Rate 0.01 Number of Cosine Annealing Epochs 100 Final Learning Rate 0.0001 Training-Validation Set Ratio 9:1 Batch Size 32

[0102] The dangerous action recognizer is used to improve the safety of the vehicle. It predicts the degree of danger based on the actions and environmental observation information output by the reinforcement learning model, and takes emergency avoidance, emergency adjustment and other actions according to the degree of danger. The dangerous action recognizer includes two stages: the sample collection stage and the training stage. In the sample collection stage, a large number of safe samples and dangerous samples are collected during the process of gradually training the strategy in the simulator environment, and the sample labels are calculated to prepare for the training of the dangerous action recognizer. Samples need to be collected in multiple scenarios and under different policy effects, and the balance of the number of positive and negative samples should be satisfied. In the training stage, the collected sample data is used for training until the end.

[0103] Refer to Figure 4 As shown, the rule constraint in step 5 is written based on human knowledge and empirical statistics of simulation experiments, and is used to restrict the behavior of the host vehicle in certain specific situations. The rule constraint mainly includes the knowledge rules of the closest distance protection, the knowledge rules under intersections, the knowledge rules before sharp turns, and the correction rules for long-term stay when there is no neighboring vehicle for the host vehicle. It judges different scenarios according to the environmental observation data and decides the speed limit rules for the host vehicle. The following is an example of the rule constraint.

[0104] In the rule constraint, the first part is mainly the constraint based on the knowledge rules of the closest distance protection: First, select the 3 environmental vehicles closest to the host vehicle. When the distance between the environmental vehicle and the host vehicle is less than 20 meters, enter the TDTC condition judgment. Judge the degree of danger according to different TDTCs and the speeds of the environmental vehicles, and then limit the maximum speed of the host vehicle. The specific rules are as follows:

[0105] 1. -2 < TDTC < 1, d / v cv <1, v cv ≥30: Limit the maximum speed of the host vehicle to 5, where d is the distance between the host vehicle and the conflicting vehicle, and v cv is the speed of the conflicting vehicle;

[0106] 2.0 < TDTC < 1.2: Limit the maximum speed of the host vehicle to 2;

[0107] 3. -1.2 < TDTC < 0: Limit the maximum speed of the host vehicle to 0.1 × planned speed;

[0108] 4. 1.2 ≤ TDTC < 3: Limit the maximum speed of the host vehicle to 15 + 20(TDTC - 1.2) / 1.8;

[0109] 5. -1.8 < TDTC ≤ -1.2, v cv > 20: Limit the maximum speed of the host vehicle to 0.1 + 1.4(TDTC - 1.2) / 0.6;

[0110] 6. -3 < TDTC ≤ -1.8, v cv > 20: Limit the maximum speed of the host vehicle to 1.4 + 6(TDTC - 1.8) / 1.2;

[0111] 7. -3 < TDTC ≤ -1.2, v cv ≤ 20: Limit the maximum speed of the host vehicle to 5 + 15(TDTC - 1.2) / 1.8;

[0112] 8. 3 ≤ TDTC < 7: Limit the maximum speed of the host vehicle to 30 + 20(TDTC - 3) / 4;

[0113] 9. -7 < TDTC ≤ 3: Limit the maximum speed of the host vehicle to 30 + 20(TDTC - 3) / 4;

[0114] 10. d < 4, v cv > 5: Limit the maximum speed of the host vehicle to 0;

[0115] 11. d < 5, v cv > 13: Limit the maximum speed of the host vehicle to 2;

[0116] 12. d < 7, v cv > 20: Limit the maximum speed of the host vehicle to 5.

[0117] The second part of the rule constraint is the knowledge rule under intersections: This rule is for collision avoidance of vehicles moving along the side direction when the host vehicle passes through an intersection. By setting rectangular judgment areas of different sizes in front of the host vehicle, it detects whether there are environmental vehicles in each area, and restricts the speed of the host vehicle by different amounts according to the environmental vehicles detected in areas of different sizes. In addition, the size of the rectangular area should be corrected according to the turning of the host vehicle, and the turning state of the host vehicle can be obtained according to the deviation angle of the host vehicle from the front road point. The specific rules are as follows:

[0118] 1. There is a non - stationary environmental vehicle within the range of 5.5 in the longitudinal distance and 10 in the lateral distance on both sides in front of the host vehicle, and this environmental vehicle satisfies TDTC < 8, vcv <4 of them: Limit the maximum speed of the host vehicle to 0;

[0119] 2. There is a non - stationary environmental vehicle within the longitudinal distance of 7 in front of the host vehicle and the lateral distances of 10 on both sides, and this environmental vehicle satisfies TDTC < 8, v cv <10 of them: Limit the maximum speed of the host vehicle to 5;

[0120] 3. There is a non - stationary environmental vehicle within the longitudinal distance of 9 in front of the host vehicle and the lateral distances of 10 on both sides, and this environmental vehicle satisfies TDTC < 8, v cv <12 of them: Limit the maximum speed of the host vehicle to 7;

[0121] 4. When the variance of the orientation angle between the host vehicle and the road point ahead is greater than 0.16, that is, when the host vehicle is turning, there is a non - stationary environmental vehicle within the longitudinal distance of 9 in front of the host vehicle and the lateral distances of 9 on both sides, and this environmental vehicle satisfies TDTC < 8, v cv <4 of them: Limit the maximum speed of the host vehicle to 0;

[0122] 5. When the variance of the orientation angle between the host vehicle and the road point ahead is greater than 0.16, that is, when the host vehicle is turning, there is a non - stationary environmental vehicle within the longitudinal distance of 10.5 in front of the host vehicle and the lateral distances of 9 on both sides, and this environmental vehicle satisfies TDTC < 8, v cv <4 of them: Limit the maximum speed of the host vehicle to 7.

[0123] The third part of the rule constraint is the correction rule for long - term stay when there is no neighboring vehicle for the host vehicle. The rule is:

[0124] 1. There is no environmental vehicle within the range of 30 of the host vehicle, and all conflicting vehicles satisfy v cv < 5, |TDTC| > 9: Manually accelerate the host vehicle and set the minimum throttle control amount to 0.3.

[0125] The fourth part of the rule constraint is the knowledge rule before a sharp turn. Speed limit is carried out according to the variance V of the orientation angle between the host vehicle and the road point ahead h and the rules are as follows:

[0126] 1. V h > 0.33: Limit the maximum speed of the host vehicle to 8;

[0127] 2. 0.25 < V h ≤0.33: Limit the maximum speed of the host vehicle to 8 + 10(0.3 - V h ) / 0.08;

[0128] 3. 0.2 < V h ≤0.25: Limit the maximum speed of the host vehicle to 14 + 10(0.25 - V h ) / 0.05;

[0129] 4.0.15 < V h ≤ 0.2: The maximum speed of the host vehicle is restricted to 22 + 20(0.2 - V h ) / 0.05;

[0130] 5.0.1 < V h ≤ 0.15: The maximum speed of the host vehicle is restricted to 30 + 20(0.15 - V h ) / 0.05;

[0131] 6.0.03 < V h ≤ 0.1: The maximum speed of the host vehicle is restricted to 40 + 20(0.1 - V h ) / 0.07;

[0132] 7.0.01 < V h ≤ 0.03: The maximum speed of the host vehicle is restricted to 60 + 20(0.03 - V h ) / 0.02;

[0133] The main function of the above rule constraint is to determine whether the environment is a dangerous situation and limit the magnitude of the actions output by reinforcement learning, which is mainly reflected in the restriction of the planned driving speed. It is reflected as the braking behavior of the vehicle in the control part. At the same time, it also performs auxiliary acceleration when the front is safe and the speed output by the reinforcement learning model is too low, ensuring the driving safety and efficiency of the vehicle.

[0134] In summary, the present invention provides a solution for safe reinforcement learning in multi-objective scenarios. This technology can be applied to fields such as intelligent vehicle assisted driving and driverless driving. Compared with traditional fully end-to-end solutions and fully rule-based solutions, it provides a new idea of a hybrid solution, combining the advantages of both to achieve the purpose of high safety, high intelligence, and high efficiency driving of the vehicle in complex and multiple scenarios. Therefore, this technology has high promotion value.

[0135] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.

Claims

1. A multi-objective complex traffic scenario-based autonomous driving solution method based on reinforcement learning, characterized in that, It includes the following steps: Step 1: Prepare a simulator environment and complex driving scenarios for autonomous driving simulation; Step 2: Add environmental feature information required for training the reinforcement learning model as environmental observation information in the observation space, including ego-vehicle information, other-vehicle information, and road information. Calculate key feature information according to the environmental feature information, including the time to collision with the vehicle in front on each lane, the time difference of the ego-vehicle's collision time from the collision point, and the variance of the orientation angle between the ego-vehicle and the front road point; Step 3: Set the reward framework required for training the reinforcement learning model; Step 4: Train the reinforcement learning model based on the time-varying training method, and then continue to train it using the meta-model in different traffic scenarios. Modify the reward weights during training according to the driving performance of the agent and the types of collisions that occur every certain number of iteration rounds, and repeat modifying the weights multiple times to end the training; Step 5: After outputting the trained reinforcement learning model, construct a dangerous action recognizer and a rule constraint device. Judge the danger level of the scenario according to the environmental observation information to limit or adjust the planned quantity output by the reinforcement learning model, and continuously manually add and optimize rules by observing the effects in the simulation environment.

2. The automatic driving solution based on reinforcement learning in a multi-objective complex traffic scenario according to claim 1, wherein In step 2, add environmental feature information required for training the reinforcement learning model in the observation space. The ego-vehicle information includes the ego-vehicle speed, the steering wheel angle of the ego-vehicle, the lane number where the ego-vehicle is located, the distance between the center point of the ego-vehicle and the center line of the lane where it is located, the deviation of the orientation angle between the fifteen road points in front of the selected ego-vehicle position and the ego-vehicle, and the relative position of the preview point at a distance of 4 / 8 meters; The other-vehicle information includes the distance of the neighboring vehicle in each lane, the time to collision with the vehicle in front on each lane, and the relative speed between the nearest neighboring other-vehicle and the ego-vehicle in the ego-vehicle lane; The road information includes the lateral distance of the ego-vehicle from the center of the lane and the orientation error of the road point.

3. The autonomous driving solution based on reinforcement learning for multi-objective complex traffic scenarios according to claim 2, wherein: The reward framework in step 3 includes an environmental reward, a speed reward, a collision penalty, and a lane center deviation penalty; among them, the environmental reward is the survival time of the ego-vehicle, which represents the duration from the starting point to the collision of the ego-vehicle; the speed reward is the driving speed of the ego-vehicle, in units of the distance traveled per second; the collision penalty is given when the ego-vehicle drives out of the route, the ego-vehicle drives out of the road boundary, or the ego-vehicle collides with an environmental vehicle; the lane center deviation penalty is the absolute value of the distance between the vehicle center and the center line.

4. A method for autonomous driving in a multi-objective complex traffic scenario based on reinforcement learning according to claim 3, characterized in that, The specific steps for training the reinforcement learning model based on the time-varying training method and combining the meta-model in step 4 are as follows: Step 4.1: Initialize the reinforcement learning model, and use each scenario to train for a certain number of rounds in turn to obtain a meta-model; Step 4.2: Use the time-varying training method to train the meta-model obtained in step 4.1 in the selected scenario, and adjust the reward weights according to the defects existing in the agent's behavior; Step 4.3: Set the scenario to all simple scenarios without intersections, and repeat the process of step 4.2 for training to improve the performance in simple scenarios without intersections; Step 4.4: Set the scenario to all scenarios with intersections, and repeat the process of step 4.2 for training to improve the performance in scenarios with intersections; Step 4.5, set the scenario to a scenario with a roundabout and multi-directional vehicles, repeat the process of Step 4.2 for training, and improve the performance in the scenario of a roundabout and multi-directional vehicles; Step 4.6, continue training in the remaining scenarios until the process ends.

5. The autonomous driving solution method in a multi-objective complex traffic scenario based on reinforcement learning according to claim 4, characterized in that, The specific steps for training the reinforcement learning model based on the time-varying training method described in Step 4 are as follows: Step 4.2.1, set the hyperparameters of the reinforcement learning model; Step 4.2.2, set the reward function as the basic reward, enable the Agent to learn lane keeping, and start iterative training; Step 4.2.3, increase the weights of the lane center deviation penalty and the collision penalty, and continue iterative training; Step 4.2.4, continue to increase the collision penalty and continue iterative training; Step 4.2.5, add new scenarios on the basis of the original scenario dataset, add speed rewards, and increase the weights of the lane center deviation penalty and the collision penalty until the iteration ends.

6. The automatic driving solution based on reinforcement learning for multi-objective complex traffic scenarios according to claim 5, wherein The dangerous action recognizer described in Step 5 predicts its danger level based on the actions and environmental observation information output by the reinforcement learning model, and takes emergency avoidance and emergency adjustment behaviors according to the danger level. The dangerous action recognizer includes a sample collection stage and a training stage. The specific steps of the sample collection stage are as follows: Step 5.1.1, prepare various types of scenarios and select a scenario to start training; Step 5.1.2, initialize the PPO policy model and start training the policy model in the selected scenario; Step 5.1.3, record the trajectory during this round of operation during the running process; Step 5.1.4, when a collision occurs, collect the 10 steps before the collision as negative samples, and randomly collect any continuous 10 steps in the trajectory of this round as positive samples; Step 5.1.5, after training until the number of running steps reaches the set total number of steps, select the next scenario to repeat the training; Step 5.1.6, end until all scenarios are collected.

7. A method for autonomous driving in a multi-objective complex traffic scenario based on reinforcement learning according to claim 6, characterized in that The dangerous action recognizer described in Step 5 is a dangerous action recognizer model based on a long short-term memory network. The specific steps of the training stage are as follows: Step 5.2.1, according to each group of sample data collected, use a sliding window to generate data groups as the model input, and use the label of the last step of each group as the target label of the model; Step 5.2.2, use the Adam optimizer and adjust the learning rate of the optimizer based on the cosine annealing method; Step 5.2.3, use the mean square error loss as the training loss function, and calculate the mean square error loss between the model output and the target label; Step 5.2.4, set the relevant model parameters and the number of training model rounds to complete the training.

8. A method for solving autonomous driving in a multi-objective complex traffic scenario based on reinforcement learning according to claim 7, characterized in that, The rule constraint in Step 5 is to write rules based on human knowledge and the empirical statistics of simulation experiments, which are used to restrict the behavior of the ego vehicle in certain specific situations. The rule constraint mainly includes the knowledge rules of the nearest distance protection, the knowledge rules under intersections, the knowledge rules before sharp turns, and the correction rules for long-term residence when there are no neighboring vehicles around the ego vehicle. The environmental observation information judges different scenarios respectively to determine the speed limit rules of the ego vehicle.

Citation Information

Patent Citations

  • Artificial intelligence automatic driving system with reinforcement learning and information fusion

    CN110764507A

  • Automatic driving decision-making method, device and equipment and storage medium

    CN114261400A