Automatic driving decision planning model training method and device

By training deep multi-agent reinforcement learning models in autonomous driving decision planning, combining instant rewards and local reward rules, and using transfer learning rules, the generalization and safety problems of deep reinforcement learning models in autonomous driving are solved, and more efficient and safe driving decisions are achieved.

CN120236163APending Publication Date: 2025-07-01WUHAN UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510283830.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing deep reinforcement learning model is insufficient in generalization and safety in autonomous driving decision planning, and it is difficult to effectively apply in complex and changeable practical driving scenarios.

Method used

By training the first deep multi-agent reinforcement learning model based on the data and preset reward rules of the first driving scenario, a second deep multi-agent reinforcement learning model is generated, and further trained in the second driving scenario using transfer learning rules to generate a target deep multi-agent reinforcement learning model. The reward rules include instant rewards and local rewards to guide vehicle decision-making behavior.

Benefits of technology

It improves the decision safety and generalization capabilities of the model in complex driving scenarios, enhances traffic efficiency and safety, and reduces the challenges of data scarcity and labeling difficulties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236163A_ABST
    Figure CN120236163A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic driving decision planning model training method and device, and belongs to the technical field of automatic driving, and the method comprises the steps: carrying out the training of a first deep multi-agent reinforcement learning model based on the data of a first driving scene and a preset reward rule; wherein the reward rule comprises the steps that when the first deep multi-agent reinforcement learning model predicts and outputs a first decision behavior of at least one vehicle according to the data of the first driving scene, an instant reward is generated for each first decision behavior, and a local reward is generated for each first vehicle set; the first vehicle set is a set of the own vehicle and the adjacent vehicle; and based on the data of the second driving scene and a preset transfer learning rule, training the first deep multi-agent reinforcement learning model again to obtain a target deep multi-agent reinforcement learning model. The target depth multi-agent reinforcement learning model obtained through training has better generalization for different driving conditions and environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to a method for training an autonomous driving decision-making and planning model. Background Art

[0002] As a core component of modern intelligent transportation systems, autonomous driving technology has made remarkable progress in recent years. However, achieving fully autonomous driving still faces many challenges, especially in complex and changing real driving scenarios.

[0003] Deep Reinforcement Learning (DRL) combines the advantages of deep learning and reinforcement learning, and can automate complex tasks by trial and error and optimizing strategies in a high-dimensional state space. DRL has demonstrated powerful capabilities in fields such as games and robot control, but its application in the decision-making and planning scenarios of autonomous driving still faces challenges in terms of generalization and safety. Summary of the Invention

[0004] In view of this, it is necessary to provide a method for training an autonomous driving decision-making and planning model to solve the problems of insufficient generalization and safety of existing deep reinforcement learning models when used for decision-making and planning of autonomous driving.

[0005] To solve the above problems, in a first aspect, the present invention provides a method for training an autonomous driving decision-making and planning model, including: Training a first deep multi-agent reinforcement learning model based on the data of a first driving scenario and a preset reward rule to obtain a second deep multi-agent reinforcement learning model; wherein, the reward rule includes that when the first deep multi-agent reinforcement learning model predicts and outputs the first decision-making behaviors of at least one vehicle according to the data of the first driving scenario, generating an immediate reward for each of the first decision-making behaviors and a local reward for a first vehicle set; the first vehicle set is a set of the ego vehicle and adjacent vehicles; Training the second deep multi-agent reinforcement learning model based on the data of a second driving scenario associated with the first driving scenario and a preset transfer learning rule to obtain a target deep multi-agent reinforcement learning model.

[0006] In a possible implementation manner, training the first deep multi-agent reinforcement learning model based on the data of a first driving scenario and a preset reward rule includes: Generating, by the first deep multi-agent reinforcement learning model, the first decision-making behaviors of each vehicle in each time step in a first simulation environment based on the first driving scenario; Generate a first immediate reward for each of the first decision-making behaviors at each time step and a first local reward for each first vehicle set at each time step; wherein, the first vehicle set is a set of the ego vehicle and adjacent vehicles in a first simulation environment. Train the first deep multi-agent reinforcement learning model according to the first immediate reward and the first local reward.

[0007] In a possible implementation, the first driving scenario is a ramp merging scenario; the first immediate reward is calculated by the following formula:

[0008] Wherein, is the first immediate reward, , , , are the weights corresponding to the collision evaluation index , the speed evaluation index , the time headway evaluation index , and the cost evaluation index for entering the lane respectively. The first local reward is calculated by the following formula:

[0009] represents the first local reward, represents the i th first vehicle set; In a possible implementation, ; , is the vehicle speed at the current time step, and are the minimum speed limit and the maximum speed limit of the vehicle respectively; , is the headway distance, is the time headway threshold; , x is the driving distance of the vehicle, L is the road length.

[0010] In a possible implementation, training the first deep multi-agent reinforcement learning model includes: When training the first deep multi-agent reinforcement learning model, determine the overall loss value of the first deep multi-agent reinforcement learning model by the following formula:

[0011] Wherein, is the overall loss value, is the policy gradient loss value, is the value function loss value, is the entropy regularization term, is the decision-making behavior in the state under the action distribution, and respectively represent and the weights of the entropy regularization term; When the overall loss value meets the preset requirements, it is determined that the training of the first deep multi-agent reinforcement learning model is completed.

[0012] In a possible implementation, the preset transfer learning rule includes: with probability select the action recommended by the second deep multi-agent reinforcement learning model before retraining, , is the initial transfer belief, is the transfer period, t is the current time step; with probability select a random behavior, , gradually decreases; with probability select the best action from the second deep multi-agent reinforcement learning model during retraining, .

[0013] In a possible implementation, the second driving scenario is an unprotected left turn scenario at an intersection; based on the data of the second driving scenario associated with the first driving scenario and the preset transfer learning rule, retraining the second deep multi-agent reinforcement learning model includes: Through the second deep multi-agent reinforcement learning model, based on the second simulation environment of the second driving scenario and the preset transfer learning rule, generate the second decision-making behavior of each vehicle in the second simulation environment at each time step; Calculate the second immediate reward for the second decision-making behavior at each time step and the second local reward for each second vehicle set at each time step through the following formula: The second vehicle set is the set of the ego vehicle and adjacent vehicles in the second simulation environment;

[0014] where, is the second immediate reward, , , are the weights corresponding to the collision evaluation index , the speed evaluation index , and the entering lane cost evaluation index respectively;

[0015] ' represents the second partial reward, indicating the i th set of second vehicles.

[0016] Train a second deep multi-agent reinforcement learning model according to the second immediate reward and the second partial reward.

[0017] In a possible implementation, greater than .

[0018] In a possible implementation, greater than , and ; greater than and .

[0019] In a second aspect, the present invention also provides an autonomous driving decision-making and planning model training device, including: A first training module for training a first deep multi-agent reinforcement learning model based on data of a first driving scenario and a preset reward rule to obtain a second deep multi-agent reinforcement learning model; wherein, the reward rule includes generating an immediate reward for each of the first decision-making behaviors and generating a partial reward for a first vehicle set when the first deep multi-agent reinforcement learning model predicts and outputs a first decision-making behavior of at least one vehicle according to the data of the first driving scenario; the first vehicle set is a set of a self-vehicle and adjacent vehicles; A second training module for training the second deep multi-agent reinforcement learning model based on data of a second driving scenario associated with the first driving scenario and a preset transfer learning rule to obtain a target deep multi-agent reinforcement learning model.

[0020] The beneficial effects of the present invention are: The present invention regards each vehicle as an agent, and then trains a first deep multi-agent reinforcement learning model based on the data of the first driving scenario and a preset reward rule. The reward rule includes that when the first deep multi-agent reinforcement learning model predicts and outputs first decision behaviors of at least one vehicle according to the data of the first driving scenario, an immediate reward is generated for each first decision behavior and a local reward is generated for the first vehicle set; the first vehicle set is a set of the ego vehicle and adjacent vehicles. The immediate reward can guide the first deep multi-agent reinforcement model to learn preferred decision behaviors from the first decision behaviors, and the local reward can guide the first deep multi-agent reinforcement model to pay attention to the interaction between the vehicle and the surrounding environment. Therefore, training the first deep multi-agent reinforcement learning model based on the data of the first driving scenario and the preset reward rule can enable the first deep multi-agent reinforcement learning model to fully learn the decision behaviors of vehicles and the interaction between vehicles in the first driving scenario, and further enable the trained second deep multi-agent reinforcement learning model to better coordinate the actions of vehicles in complex driving scenarios, improving the overall traffic efficiency and safety.

[0021] Then, based on the data of a second driving scenario associated with the first driving scenario and a preset transfer learning rule, the second deep multi-agent reinforcement learning model is retrained to obtain a target deep multi-agent reinforcement learning model. By means of transfer learning, the requirement for the data of the second driving scenario can be reduced, the challenges brought by data scarcity and annotation difficulties can be reduced, the generalization ability of the second deep multi-agent reinforcement learning model to different driving conditions and environments can be improved, and its generalization ability can be enhanced, so that the obtained target deep multi-agent reinforcement learning model has higher decision-making safety and generalization ability in actual driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0023] Figure 1 It is a schematic flowchart of an embodiment of the method for training an autonomous driving decision-making and planning model provided by the present invention; Figure 2 For the present invention Figure 1 It is a schematic flowchart of an embodiment of S101 in Figure 3 It is a schematic diagram of a ramp merging environment provided by the present invention; Figure 4 For the present invention Figure 1Flow diagram of an embodiment of S102 in the present invention; Figure 5 Schematic diagram of an unprotected left-turn environment at an intersection provided by the present invention; Figure 6 Schematic diagram of a transfer learning rule provided by the present invention; Figure 7 Flow chart of training an autonomous driving decision-making and planning model provided by the present invention; Figure 8 Another flow chart of training an autonomous driving decision-making and planning model provided by the present invention; Figure 9 Schematic structural diagram of an embodiment of an apparatus for training an autonomous driving decision-making and planning model provided by the present invention. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0025] In the description of the embodiments of the present invention, unless otherwise specified, "a plurality of" means two or more. In the embodiments of the present invention, the "first", "second", etc. are used to distinguish similar objects, rather than to describe a specific order or sequence, nor to indicate or imply their relative importance or implicitly specify the quantity of the indicated technical features. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by the "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or more.

[0026] Referring to "

[0027] Referring Figure 1 , a flow diagram of an embodiment of a method for training an autonomous driving decision-making and planning model provided by the present invention is shown. The method includes: S101. Train a first deep multi-agent reinforcement learning model for driving strategy planning based on the data of the first driving scenario and a preset reward rule to obtain a second deep multi-agent reinforcement learning model. The reward rule includes generating an immediate reward for each first decision-making behavior and a local reward for a first vehicle set when the first deep multi-agent reinforcement learning model predicts and outputs the first decision-making behaviors of at least one vehicle according to the data of the first driving scenario. The first vehicle set is a set of the ego vehicle and adjacent vehicles.

[0028] The first driving scenario can be a ramp merging scenario, where ramp merging refers to the scenario of a vehicle entering the main road from a ramp. The data of the first driving scenario can include data such as the road information and vehicle information of the first driving scenario.

[0029] The reward rule can include an immediate reward rule and a local reward rule. The immediate reward rule is used to generate an immediate reward for each first decision-making behavior when the first deep multi-agent reinforcement learning model predicts and outputs the first decision-making behaviors of at least one vehicle according to the data of the first driving scenario. Specifically, collision evaluation indicators, speed evaluation indicators, time headway evaluation indicators, and lane entry cost evaluation indicators, etc., of each first decision-making behavior can be determined, and then the immediate reward of the first decision-making behavior can be generated based on these indicators. The immediate reward can guide the first deep multi-agent reinforcement model to learn preferred decision-making behaviors from the first decision-making behaviors.

[0030] The local reward rule is used to generate a local reward for the first vehicle set when the first deep multi-agent reinforcement learning model predicts and outputs the first decision-making behaviors of at least one vehicle according to the data of the first driving scenario. The first vehicle set is a set of the ego vehicle and adjacent vehicles, and each vehicle can be used as the ego vehicle, so there can be multiple first vehicle sets. Specifically, the sum of the immediate rewards of each vehicle in the first vehicle set can be used as the local reward of the first vehicle set. The local reward can guide the first deep multi-agent reinforcement model to pay attention to the mutual influence between the vehicle and the surrounding environment.

[0031] The first deep multi-agent reinforcement learning model can be a deep multi-agent reinforcement learning model for driving strategy planning. The deep multi-agent reinforcement learning model can be an independent learning model, a centralized training and decentralized execution model, a fully centralized model, etc. Specifically, it can be selected according to the actual situation.

[0032] Each vehicle in the first driving scenario can be regarded as an agent. Through the first deep multi-agent reinforcement learning model, based on the data of the first driving scenario and the preset reward rule, learn the decision-making behaviors and interaction situations of each vehicle in the first driving scenario, so as to obtain the second deep multi-agent reinforcement learning model.

[0033] S102. Re-train the second deep multi-agent reinforcement learning model based on the data of the second driving scenario associated with the first driving scenario and the preset transfer learning rules to obtain the target deep multi-agent reinforcement learning model.

[0034] The second driving scenario is a driving scenario different from but associated with the first driving scenario. For example, it can be an unprotected left turn scenario at an intersection, which means a scenario where a left-turning vehicle makes a left turn without a dedicated signal light or left-turn lane. The data of the second driving scenario can be data such as road information and vehicle information of the second driving scenario. The second deep multi-agent reinforcement learning model can be re-trained based on the preset transfer learning rules and the data of the second driving scenario to obtain the target deep multi-agent reinforcement learning model.

[0035] The target deep multi-agent reinforcement learning model obtained after the two trainings can be used in actual applications to formulate driving strategies for the second driving scenario.

[0036] In the present invention, each vehicle is regarded as an agent, and then the first deep multi-agent reinforcement learning model is trained based on the data of the first driving scenario and the preset reward rules. It can enable the first deep multi-agent reinforcement learning model to fully learn the decision-making behaviors of vehicles and the interaction situations between vehicles in the first driving scenario, so that the trained second deep multi-agent reinforcement learning model can better coordinate the actions of vehicles in complex driving scenarios and improve the overall traffic efficiency and safety.

[0037] Then, re-train the second deep multi-agent reinforcement learning model based on the data of the second driving scenario associated with the first driving scenario and the preset transfer learning rules to obtain the target deep multi-agent reinforcement learning model. Through the method of transfer learning, the demand for the data of the second driving scenario can be reduced, the challenges brought by data scarcity and annotation difficulties can be reduced, and the generalization ability of the second deep multi-agent reinforcement learning model to different driving conditions and environments can be improved, so that the obtained target deep multi-agent reinforcement learning model has higher decision-making safety and generalization ability in actual driving scenarios.

[0038] In some embodiments of the present invention, as Figure 2 shown, S101 includes: S201. Through the first deep multi-agent reinforcement learning model, generate the first decision-making behaviors of each vehicle in the first simulation environment at each time step based on the first simulation environment of the first driving scenario.

[0039] The target task of the target vehicle in the first simulation environment can be defined first. The target vehicle is an autonomous vehicle that needs to complete ramp merging, and the target task is to identify the ramp entrance, maintain a safe distance, avoid other vehicles, and smoothly enter the ramp.

[0040] Then, the first simulation environment can be constructed through the OpenAI Gym platform. OpenAI Gym is a widely used open-source library that provides a set of standard interfaces and various simulation environments for developing and comparing reinforcement learning algorithms to ensure that reinforcement learning models can be effectively trained and tested in complex traffic scenarios such as the first driving scenario and the second driving scenario.

[0041] When the first driving scenario is a ramp merging scenario, the configuration of the first simulation environment on the OpenAI Gym platform is as follows: Setting of human vehicle types: Set vehicle types such as cars, trucks, etc.

[0042] Definition of state space and action space: Define the state space and action space of the vehicle to ensure that the model can correctly perceive the environment and make appropriate decisions. In a multi-agent environment, each agent needs to understand its own state, the state of the surrounding environment, and the states of other agents, and at the same time select appropriate actions to achieve the goal.

[0043] Vehicle speed range: Limit the speed range of the vehicle to ensure that the vehicle speed conforms to the required environmental vehicle speed.

[0044] Degree of environmental complexity: Set the number of interacting vehicles to simulate traffic scenarios with different densities and complexities, thereby providing training environments with different difficulties.

[0045] In addition, the road layout, traffic lights, vehicle driving rules, etc. can also be configured. The above configuration settings will help to effectively build the simulation environment and train the reinforcement learning model on the OpenAI Gym platform.

[0046] For example, the setting standards for the ramp merging area scenario are as follows: ① Human vehicle types: Cars, using the IDM and MOBIL models.

[0047] ② Vehicle speed range: The initial speed is randomly selected between 25 - 27 m / s, and the maximum speed is 40 m / s.

[0048] ③ Degree of environmental complexity: Simple mode: 1 - 3 autonomous vehicles and 1 - 3 human vehicles; Medium mode: 2 - 4 autonomous vehicles and 2 - 4 human vehicles; Difficult mode: 4 - 6 autonomous vehicles and 4 - 6 human vehicles.

[0049] ④ Construction of ramp merging environment: Length of on-ramp straight section : 220 m; Length of ramp-in inclined section : 100 m; Length of the section connecting the ramp and the main road : 100 m; Length of the last main road terminal section of the ramp terminal : 100 m.

[0050] like Figure 3 As shown, a schematic diagram of a ramp merging environment provided by the present invention is shown. The green vehicle is a vehicle entering the main road from the ramp. x It is the distance traveled by the vehicle at the junction of the on-ramp and the main road.

[0051] ⑤Behavior space: The behavior space of agent i (target vehicle) It is defined as a set of decision-making behaviors of left, cruising, accelerating, and decelerating. The overall action space of the system is the joint action of each autonomous driving vehicle, that is, .

[0052] State Space: The agent The state space Defined as Dimension The matrix of is the number of observed vehicles, The number of features that characterize the vehicle state, including: Ispresent: A Boolean variable that determines whether the vehicle is observable near the ego vehicle.

[0053] : The longitudinal position of the observed vehicle relative to the ego vehicle.

[0054] : The lateral position of the observed vehicle relative to the ego vehicle.

[0055] : The longitudinal velocity of the observed vehicle relative to the ego vehicle.

[0056] : The lateral velocity of the observed vehicle relative to the ego vehicle.

[0057] In the present invention, only “neighboring vehicles” can be observed by the ego vehicle. “Neighboring vehicles” are defined as vehicles within 150 m of the longitudinal distance from the ego vehicle. The entire state of the environment is the Cartesian product of each state, i.e. .

[0058] After completing the above definitions and configurations, the first decision-making behaviors of each vehicle in the first simulation environment at each time step can be generated based on the target task and the first simulation environment by means of the first deep multi-agent reinforcement learning model.

[0059] S202, generate a first immediate reward for each first decision-making behavior at each time step and a first local reward for each first vehicle set at each time step; wherein, the first vehicle set is a set of the ego vehicle and adjacent vehicles in the first simulation environment.

[0060] The first immediate reward can be calculated by the following formula:

[0061] wherein, is the first immediate reward, , , , are the weights corresponding to the collision evaluation index , the speed evaluation index , the time headway evaluation index , and the cost evaluation index for entering the lane respectively. , , and can be determined according to the custom collision evaluation rules, speed evaluation, time headway evaluation, and cost evaluation rules for entering the lane.

[0062] In some embodiments, , , and can be calculated by the following rules: . , is the vehicle speed of the ego vehicle at the current time step, and are the minimum speed limit and the maximum speed limit of the ego vehicle set respectively. For example, they can be set to and respectively. , is the distance between the front of the ego vehicle and the rear of the vehicle in front, is the time headway threshold. When the time headway is less than , the ego vehicle will be penalized. Only when the time headway is greater than , the ego vehicle will be rewarded. The time headway threshold can be 1.2 s. , x is the driving distance of the ego vehicle, and L is the length of the road. When the ego vehicle is a vehicle that needs to merge into the ramp,x It can be the driving distance of the ego vehicle at the connection section between the ramp and the main road, L is the length of the connection section between the ramp and the main road, that is is to penalize the waiting time on the lane to avoid deadlocks.

[0063] Since safety and anti-collision are the most important criteria, it can make much larger than other weights to prioritize safety. For example, set , , , .

[0064] The first partial reward can be calculated by the following formula:

[0065] represents the first partial reward, represents the i th set of first vehicles.

[0066] This design of the partial reward enables the model to focus only on the agents most relevant to the success or failure of the task.

[0067] S203. Train the first deep multi-agent reinforcement learning model according to the first immediate reward and the first partial reward.

[0068] The first immediate reward and the first partial reward can guide the first deep multi-agent reinforcement learning model to learn the optimal strategy in the first driving scenario. The immediate reward is given according to the behavior of each agent at each time step, such as avoiding collisions, staying in the lane, and passing obstacles. The partial reward strategy can solve the problems of delay, communication overhead, and credit assignment brought by sharing the global reward by focusing on the local environment and mutual influence of the agents. This setting enables the agents to optimize their own behaviors while considering the performance of the local environment system, thereby achieving the overall optimization of the system.

[0069] In some embodiments of the present invention, the process of training the first deep multi-agent reinforcement learning model may include: when training the first deep multi-agent reinforcement learning model, determine the overall loss value of the first deep multi-agent reinforcement learning model through the following formula:

[0070] Among them, is the overall loss value, is the policy gradient loss value, is the value function loss value, is the entropy regularization term, is the decision-making behavior The action distribution in the state , specifically, , is used to encourage the agent to explore new states; and respectively represent and the weights of the entropy regularization term, and can be 1 and 0.01 respectively; When the overall loss value meets the preset requirements, it is determined that the training of the first deep multi-agent reinforcement learning model is completed.

[0071] The policy gradient loss value algorithm can include the REINFORCE algorithm, the Actor-Critic algorithm, the PPO algorithm, etc. In the embodiments of the present invention, the specific calculation formula of the policy gradient loss value can be: , where is the expectation of the decision-making behavior , is the probability of the decision-making behavior selecting the action in the state , is the advantage function, is the state value function, is the discount factor.

[0072] The value function loss value can include the state value function loss value, the action value function loss value, the temporal difference error loss value, etc. In the embodiments of the present invention, the specific calculation formula of the value function loss value can be:

[0073] where represents the experience replay buffer for collecting previously encountered experiences, represents the parameters obtained from the previous iteration, represents the value function estimate of the next state, represents the value function estimate of the current state.

[0074] In some embodiments of the present invention, when training the first deep multi-agent reinforcement learning model, curriculum learning can be used to gradually increase the task difficulty and decompose complex tasks, so that the agent can learn and adapt to complex environments more efficiently, and improve its generalization ability and learning efficiency.

[0075] Use curriculum learning to accelerate students' learning speed and improve performance in the hard mode. Specifically, instead of directly learning the hard mode, train a well-trained model from easier modes (i.e., easy and medium), and train the model to achieve higher efficiency. Curriculum learning is particularly desirable for safety-critical tasks because starting from a suitable model can greatly reduce the number of potentially risky "blind" explorations.

[0076] In some embodiments of the present invention, as Figure 4 shown, S102 may include: S401, through the second deep multi-agent reinforcement learning model, based on the second simulation environment of the second driving scenario and the preset transfer learning rules, generate the second decision-making behaviors of each vehicle in the second simulation environment at each time step.

[0077] The target vehicle in the second simulation environment is an autonomous vehicle that needs to complete an unprotected left turn at an intersection, and the target task is to identify the intersection, judge the left-turn timing, avoid oncoming vehicles, and smoothly complete the left turn.

[0078] The second simulation environment of the second driving scenario can also be constructed through the OpenAI Gym platform. The configuration of the human vehicle type, vehicle speed range, and environmental complexity of the second simulation environment on the OpenAI Gym platform can be configured with reference to the above examples. However, due to the different driving scenarios, some configurations can be adaptively modified. For example, referring to Figure 5 , shows a schematic diagram of an unprotected left-turn environment at an intersection provided by the present invention. The construction of the unprotected left-turn environment at the intersection: the length of the two-way lane is: 220 m.

[0079] The preset transfer learning rules may include: ① Transfer rule: With probability select the action recommended by the second deep multi-agent reinforcement learning model before retraining, , is the initial transfer belief, is the transfer period, t is the current time step.

[0080] ② Exploration rule: With probability select a random behavior for exploring new strategies. This rule is complementary to the previous rule, , gradually decreases from 1 to 0.1.

[0081] ③ Exploitation rule: With probability select the best action from the second deep multi-agent reinforcement learning model during retraining, .

[0082] Under the guidance of these three transfer learning rules, the agent can effectively utilize the knowledge from the second deep multi-agent reinforcement learning model before training, namely the expert network.

[0083] Refer to Figure 6 , which shows a schematic diagram of a transfer learning rule provided by the present invention. The initial value of is , and as time increases, T tran gradually decreases and reaches 0 at the time point. As time increases, T tran also gradually decreases. At the time point, Figure 6 stops changing. After the T tran time point, it is not drawn anymore. When reaching the Tt ran time point, and both stop changing, while continues to decrease. The agent more relies on the strategy of the second deep multi-agent reinforcement learning model in the retraining to make decisions. Finally, it stabilizes at the T exp time point.

[0084] Decreasing as time increases indicates that the influence of the source domain task gradually decreases, and the control strategy of the target domain task gradually dominates. also decreases as time increases. In the early stage of training, is larger, representing more exploration. As the training progresses, decreases, and the agent's choice is more inclined to utilize the learned strategy. The change of

[0085] is based on the decrease of ε, indicating that the agent's dependence on the target task network gradually increases.

[0086] S402. Calculate the second immediate reward for the second decision-making behavior at each time step and the second local reward for each second vehicle set at each time step: The second vehicle set is the set of ego vehicle and neighboring vehicles in the second simulation environment.

[0087] The second immediate reward can be calculated by the following formula:

[0088] Where is the second immediate reward, , , are the weights corresponding to the collision evaluation index , the speed evaluation index , and the cost evaluation index for entering the lane respectively. , and The calculation formulas of can refer to the above. In the second driving scenario x in the calculation formula can be the driving distance of the vehicle on the two-way lane, L can be the length of the two-way lane, that is .

[0089] In the ramp merging scenario, there are collision evaluation index , speed evaluation index , time headway evaluation index , and cost evaluation index for entering the lane these indexes. However, in the scenario of left-turn at an unsignalized intersection, the vehicles move in an orderly manner, so the time headway evaluation index is removed, and an additional penalty term for uncertainty is added to consider the risk and delay of the left-turn action.

[0090] The scenario of left-turn at an unsignalized intersection pays attention to both safety and traffic efficiency compared with the ramp merging scenario. Therefore, , , can be set.

[0091] The second local reward can be calculated by the following formula:

[0092] ' represents the second local reward, represents the i th second vehicle set.

[0093] S403. Train the second deep multi-agent reinforcement learning model according to the second immediate reward and the second local reward.

[0094] The second immediate reward and the second local reward can guide the second deep multi-agent reinforcement learning model to learn the optimal policy in the second driving scenario.

[0095] In some embodiments of the present invention, the above formula for calculating the overall loss value can still be used to calculate the overall loss value when the second deep multi-agent reinforcement learning model is trained based on the second simulation environment. It should be noted that in the formula for calculating the value function loss value, can be adaptively modified to .

[0096] In some embodiments of the present invention, when training the second deep multi-agent reinforcement learning model, curriculum learning can also be used to gradually increase the task difficulty and decompose complex tasks.

[0097] In some embodiments of the present invention, after the training of the second deep multi-agent reinforcement learning model is completed, evaluation metrics of the model can be generated, and the learning rate of the second deep multi-agent reinforcement learning model, the weights in the reward mechanism, and the weighting coefficients in the formula for calculating the overall loss value can be optimized according to the evaluation metrics.

[0098] The evaluation metrics can include: ① Task completion assessment: The proportion of the number of times the specified task is successfully completed to the total number of attempts. This metric is used to evaluate the success rate of the autonomous driving model when performing a specific task. A high task completion rate means that the model can reliably complete the specified task, indicating good task execution ability.

[0099] ② Average reward assessment: The average reward value obtained by the model during the execution of the task. This metric is used to evaluate the success rate of the autonomous driving model when performing a specific task. A high task completion rate means that the model can reliably complete the specified task, indicating good task execution ability.

[0100] ③ Collision rate assessment: The proportion of the number of collisions that occur during the execution of the task to the total number of attempts. This metric is used to evaluate the safety of the autonomous driving model. A low collision rate means that the model can effectively avoid collisions and has high safety.

[0101] Model optimization can include: ① Adjusting the learning rate of the model , to balance the training speed and stability. The optimal learning rate is found through Bayesian optimization.

[0102]

[0103] Among them, is the objective function, the optimal learning rate.

[0104] ② Redesign and adjust the weights in the reward mechanism according to the performance of the model in the second simulation environment 、 、 to make it better reflect the desired behavior.

[0105] ③ Adjust the weighting coefficients of the value function loss and the entropy regularization term and to optimize the comprehensive loss function.

[0106] Referring to Figure 7 shows a training flowchart of an autonomous driving decision-making and planning model provided by the present invention. First, create the required environments, including creating a first simulation environment and a second simulation environment, and defining the target tasks of the target vehicle in the first simulation environment and the second simulation environment (source domain task and target domain task). Secondly, build the model, including defining the state space and the behavior space, designing the reward mechanism, defining the loss function, and using curriculum learning to decompose the difficulty. Thirdly, customize the transfer learning rules, including transfer rules, exploration rules, and exploitation rules. Finally, evaluate and improve, including redesigning the loss function, evaluation metrics, and analyzing improvements.

[0107] Referring to Figure 8 shows another training flowchart of an autonomous driving decision-making and planning model provided by the present invention. Build a ramp merge simulation environment and an unprotected intersection left-turn simulation environment through the OpenAI Gym platform, train a first deep multi-agent reinforcement learning model based on the ramp merge simulation environment, and train a second deep multi-agent reinforcement learning model based on the unprotected intersection left-turn simulation environment.

[0108] Referring to Figure 9 shows a schematic structural diagram of an embodiment of an apparatus for training an autonomous driving decision-making and planning model provided by the present invention. The apparatus 900 includes: A first training module 901, configured to train a first deep multi-agent reinforcement learning model based on data of a first driving scenario and a preset reward rule to obtain a second deep multi-agent reinforcement learning model; wherein, the reward rule includes that when the first deep multi-agent reinforcement learning model predicts and outputs a first decision behavior of at least one vehicle according to the data of the first driving scenario, an immediate reward is generated for each of the first decision behaviors and a local reward is generated for a first vehicle set; the first vehicle set is a set of ego vehicle and adjacent vehicles; A second training module 902, configured to train the second deep multi-agent reinforcement learning model based on data of a second driving scenario associated with the first driving scenario and a preset transfer learning rule to obtain a target deep multi-agent reinforcement learning model.

[0109] It should be noted that: The implementation principles or implementation processes of the above-mentioned modules can refer to the embodiments of the aforementioned method for training an autonomous driving decision-making and planning model, and will not be elaborated one by one here.

[0110] Those skilled in the art can understand that all or part of the processes for implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory, or a random access memory, etc.

[0111] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for training an autonomous driving decision planning model, characterized in that: include: Based on the data of the first driving scene and the preset reward rule, the first deep multi-agent reinforcement learning model is trained to obtain a second deep multi-agent reinforcement learning model; wherein the reward rule includes generating an immediate reward for each first decision-making behavior and generating a local reward for a first vehicle set when the first deep multi-agent reinforcement learning model predicts and outputs a first decision-making behavior of at least one vehicle according to the data of the first driving scene; the first vehicle set is a set of the self vehicle and adjacent vehicles; Based on data of a second driving scenario associated with the first driving scenario and a preset transfer learning rule, the second deep multi-agent reinforcement learning model is trained to obtain a target deep multi-agent reinforcement learning model.

2. The autonomous driving decision planning model training method according to claim 1, characterized in that: Based on the data of the first driving scenario and the preset reward rules, the first deep multi-agent reinforcement learning model is trained, including: Generate, by a first deep multi-agent reinforcement learning model, a first decision-making behavior of each vehicle in the first simulation environment at each time step based on a first simulation environment of a first driving scenario; Generate a first immediate reward for each of the first decision behaviors at each time step and generate a first local reward for each of the first vehicle sets at each time step; wherein the first vehicle set is a set of the self vehicle and adjacent vehicles in the first simulation environment; The first deep multi-agent reinforcement learning model is trained according to the first immediate reward and the first local reward.

3. The autonomous driving decision planning model training method according to claim 2, characterized in that: The first driving scenario is a ramp merging scenario; the first instant reward is calculated using the following formula: in, For the first instant reward, , , , Collision evaluation index , speed evaluation index , Headway evaluation index , Lane entry cost evaluation index The corresponding weight; The first partial reward is calculated by the following formula: represents the first partial reward, Indicates i The first vehicle collection.

4. The automatic driving decision-making method according to claim 3, characterized in that: ; , is the vehicle speed at the current time step, and are the minimum and maximum speed limits of the vehicle respectively; , is the headway distance, is the headway threshold; , x is the vehicle travel distance, L is the length of the road.

5. The autonomous driving decision planning model training method according to claim 1, characterized in that: Training the first deep multi-agent reinforcement learning model, including: When training the first deep multi-agent reinforcement learning model, the overall loss value of the first deep multi-agent reinforcement learning model is determined by the following formula: in, is the overall loss value, is the policy gradient loss value, is the loss value of the value function, is the entropy regularization term, For decision making In Status The action distribution under and Respectively and the weight of the entropy regularization term; When the overall loss value meets the preset requirements, it is determined that the training of the first deep multi-agent reinforcement learning model is completed.

6. The autonomous driving decision planning model training method according to claim 1, characterized in that: The preset transfer learning rules include: Select the action suggested by the second pre-trained deep multi-agent reinforcement learning model, , is the initial migration belief, For the migration cycle, t is the current time step; with probability Select random behavior, , Gradually reduce; with probability Select the best action from the second deep multi-agent reinforcement learning model in training, .

7. The autonomous driving decision planning model training method according to claim 3, characterized in that: The second driving scenario is an unprotected left turn scenario at an intersection; based on data of the second driving scenario associated with the first driving scenario and a preset transfer learning rule, the second deep multi-agent reinforcement learning model is trained, including: Generate, by the second deep multi-agent reinforcement learning model, a second decision-making behavior of each vehicle in the second simulation environment at each time step based on a second simulation environment of a second driving scenario and a preset transfer learning rule; The second immediate reward for the second decision behavior at each time step and the second local reward for each second vehicle set at each time step are calculated by the following formula: the second vehicle set is a set of the self vehicle and the adjacent vehicles in the second simulation environment; in, For the second instant reward, , , Collision evaluation index , speed evaluation index , Lane entry cost evaluation index The corresponding weight; ' indicates the second partial reward, Indicates i A second vehicle set; According to the second immediate reward and the second local reward, a second deep multi-agent reinforcement learning model is trained.

8. The autonomous driving decision planning model training method according to claim 7, characterized in that: Greater than .

9. The autonomous driving decision planning model training method according to claim 7, characterized in that: Greater than , and ; Greater than and .

10. An automatic driving decision planning model training device, characterized in that: include: A first training module is used to train the first deep multi-agent reinforcement learning model based on the data of the first driving scene and a preset reward rule to obtain a second deep multi-agent reinforcement learning model; wherein the reward rule includes generating an immediate reward for each first decision-making behavior and generating a local reward for a first vehicle set when the first deep multi-agent reinforcement learning model predicts and outputs a first decision-making behavior of at least one vehicle according to the data of the first driving scene; the first vehicle set is a set of the self vehicle and adjacent vehicles; The second training module is used to train the second deep multi-agent reinforcement learning model based on data of a second driving scenario associated with the first driving scenario and a preset transfer learning rule to obtain a target deep multi-agent reinforcement learning model.

Citation Information

Cited By

  • Driving strategy model training method, automatic driving method and related device

    CN120821268A

  • Vehicle driving track planning model training method, vehicle driving track planning method and electronic equipment

    CN121144844A