An autonomous driving strategy generation method, device, and storage medium
By introducing the experience playback classification pool mechanism and the extraction of previous vehicle strategies into the DDPG algorithm model of autonomous driving vehicles, the problem of low learning efficiency of autonomous driving vehicles under limited time and computing power is solved, and more efficient learning and training effects are achieved, meeting the real-time requirements of autonomous driving.
Patent Information
- Application Number
- CN202210328209.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-03-31
AI Technical Summary
The prior art under limited time and computing power conditions, the learning efficiency of autonomous driving vehicles is low and cannot meet the real-time requirements of automobile autonomous driving.
The autonomous driving strategy model based on the DDPG algorithm is adopted to build a simulation environment for autonomous driving, simple, ordinary and difficult training tasks are formulated, and the experience playback classification pool mechanism is introduced to conduct experience classification learning, and the previous vehicle strategy experience is extracted for learning.
It improves the learning efficiency and training effect of autonomous driving vehicles, shortens the convergence time of the model, improves the driving ability for complex scenarios, and meets the real-time requirements of autonomous driving.
Smart Images

Figure CN114771561B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, device and storage medium for generating an autonomous driving strategy, belonging to the technical field of driverless driving. Background Art
[0002] With the development of social economy, as a means of transportation, automobiles have been widely used in social production. However, with the wide application of automobiles, the number of traffic accidents has been increasing year by year. The application of artificial intelligence-based autonomous driving technology can improve the carrying capacity of highways, reduce greenhouse gas emissions, avoid potential safety hazards caused by human factors (fatigue driving, drunk driving), and reduce the accident rate, which is of great significance for the improvement of the future traffic ecosystem.
[0003] The control of autonomous vehicles is mostly based on rule control. In complex scenarios, the number of rules will increase exponentially, and conflicts may also occur between rules. Secondly, the testing and verification of autonomous vehicles under complex working conditions are often difficult to carry out due to safety considerations. Different from supervised learning, reinforcement learning can perform autonomous learning through data-driven or interaction with the environment. After sufficient training, it can autonomously handle complex working conditions, so it is more suitable for the decision-making and control of autonomous driving.
[0004] The learning effect of the agent is better in simple road conditions. However, in complex road conditions, especially when facing dangerous driving situations such as pedestrians running red lights, trucks changing lanes, foggy weather, narrow and curved sections, etc., the learning efficiency of the agent is not high, the convergence speed is slow, and it cannot meet the requirements of instantaneity and high efficiency of automobile driving. How to improve the learning efficiency of reinforcement learning and meet the real-time requirements of automobile autonomous driving under limited time and computing power conditions is a problem that needs to be solved by the existing technology. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies in the prior art, and provide a method, device and storage medium for generating an autonomous driving strategy, so as to solve the technical problems that the existing technology has low learning efficiency and cannot meet the real-time requirements of autonomous driving under limited time and computing power conditions.
[0006] To achieve the above purpose, the present invention is implemented by adopting the following technical solutions:
[0007] In the first aspect, the present invention provides a method for generating an autonomous driving strategy, including:
[0008] When obtaining an autonomous driving task, initialize the vehicle's own state;
[0009] Generate an autonomous driving strategy based on the trained autonomous driving strategy model according to the own state;
[0010] Execute the autonomous driving task according to the autonomous driving strategy;
[0011] Among them, the training process of the autonomous driving strategy model includes:
[0012] Construct a simulation environment for autonomous driving, and formulate simple training tasks, ordinary training tasks, and difficult training tasks according to the simulation environment;
[0013] Construct an autonomous driving strategy model based on the DDPG algorithm;
[0014] Initialize the autonomous driving strategy model and conduct one training based on the simple training task;
[0015] Conduct secondary training on the autonomous driving strategy model after one training based on the ordinary training task;
[0016] Conduct tertiary training on the autonomous driving strategy model after secondary training based on the difficult training task.
[0017] Optionally, the simple training task is a single vehicle driving in clear weather, the ordinary training task is a single vehicle driving in foggy weather, and the difficult training task is multiple vehicles driving in foggy weather.
[0018] Optionally, the initialization of the autonomous driving strategy model and the one-time training based on the simple training task include:
[0019] Initialize the autonomous driving strategy model, including initializing the real action network, target action network, real evaluation network, and target evaluation network of the autonomous driving strategy model;
[0020] Obtain a model training sample set by executing the simple training task based on the autonomous driving strategy model in the simulation environment of autonomous driving;
[0021] Train and update the autonomous driving strategy model through the model training sample set, and replace the initialized autonomous driving strategy model with the updated autonomous driving strategy model and bring it into the above steps for iteration;
[0022] If the preset maximum number of iterations is reached, one training is completed.
[0023] Optionally, the obtaining of the model training sample set includes:
[0024] Obtain the current state S of the vehicle under the simple training task, and bring it into the initialized autonomous driving strategy model π to obtain the current action strategy A of the vehicle:
[0025] A = π(φ(S)) + N
[0026] Among them, φ(S) is the feature vector of state S, π(·) is the action policy generated by the real-world action network, and N is the random noise function;
[0027] Drive the vehicle to autopilot in the simulation environment according to the current action policy A and obtain the next state S of the vehicle * , and according to the next state S of the vehicle * Obtain the action score R and the termination state E;
[0028] Save the feature vector φ(S) of the current state of the vehicle, the feature vector φ(S * ), the current action policy A, the training score R, and the termination state E as an experience replay array, denoted as {φ(S), φ(S * ), A, R, E};
[0029] According to the next state S of the vehicle * Judge whether there is a vehicle in front. If so, obtain the experience replay array of the vehicle in front and store it in the pre-constructed experience replay set D of the vehicle in front M ;
[0030] Judge whether the vehicle terminates according to the termination state. If not, store the experience replay array in the pre-constructed successful experience replay set D S , and bring the next state of the vehicle into the above steps to execute in a loop;
[0031] If so, store the experience replay array in the pre-constructed failed experience replay set D F , and output the failed experience replay set D F , the successful experience replay set D S and the experience replay set D of the vehicle in front M ;
[0032] Randomly select T experience replay arrays from the successful experience replay set D S , the failed experience replay set D F and the experience replay set D of the vehicle in front M to generate a model training sample set D T .
[0033] Optionally, the state includes vehicle coordinates, speed, acceleration, and vehicle surrounding environment information. The vehicle surrounding environment information includes obstacles within a preset range around the vehicle coordinates. The obstacles include vehicles in front, fences, and signal lights.
[0034] Optionally, the obtaining of the training score R and the termination state E according to the next state S of the vehicle * includes:
[0035] According to the current state S and the next state S *The vehicle coordinates in it obtain the driving trajectory and driving distance of the vehicle;
[0036] Obtain the training base score according to the driving distance and the preset unit distance score;
[0037] Judge whether the vehicle collides with an obstacle according to the driving trajectory and the obstacle coordinates. If a collision occurs, deduct the preset collision score from the training base score and set the termination state E to terminated;
[0038] Judge whether the vehicle passes the traffic light according to the driving trajectory and the traffic light coordinates. If it passes, calculate the time when the vehicle arrives at the traffic light according to the speed and acceleration of the current state, and judge the state of the traffic light according to the arrival time of the traffic light. If the traffic light is red, deduct the preset running red light score from the training base score;
[0039] Obtain the deducted training base score as the training score R.
[0040] Optionally, training and updating the autonomous driving policy model through the model training sample set includes:
[0041] According to the model training sample set D T Calculate the current target reward value y t :
[0042]
[0043] where y t is the target reward value of the t-th experience replay array, t = 1, 2, 3... T, R t is the training score of the t-th experience replay array, γ is the discount factor, π′(·) is the action policy generated by the target action network, ω′ is the weight parameter of the target evaluation network, Q′(·) is the evaluation value generated by the target evaluation network, is the current state S t of the t-th experience replay array;
[0044] Based on the target reward value y t Construct the first loss function, and update the weight parameter ω of the real evaluation network through the gradient backpropagation of the neural network; the first loss function is:
[0045]
[0046] where ω is the weight parameter of the real evaluation network, Q(·) is the evaluation value generated by the real target evaluation network; A t is the current action policy of the t-th experience replay array;
[0047] Construct a second loss function, and update the weight parameter θ of the real action network through the gradient backpropagation of the neural network; the second loss function is as follows:
[0048]
[0049] If t%C = 0, then update the weight parameters ω' and θ' of the target action network and the target evaluation network according to the weight parameters ω and θ of the real action network and the real evaluation network:
[0050] ω' ← τω + (1 - τ)ω'
[0051] θ' ← τθ + (1 - τ)θ'
[0052] where C is the update frequency of the weight parameters of the target action network and the target evaluation network, and τ is the update coefficient;
[0053] Update the autonomous driving policy model based on the updated weight parameters ω, θ, ω' and θ' of the real action network, the real evaluation network, the target action network and the target evaluation network.
[0054] In a second aspect, the present invention provides an autonomous driving policy generation device, the device includes:
[0055] An initialization module, configured to initialize the vehicle's own state when obtaining an autonomous driving task;
[0056] A policy generation module, configured to generate an autonomous driving policy based on the trained autonomous driving policy model according to the own state;
[0057] An autonomous driving module, configured to execute an autonomous driving task according to the autonomous driving policy;
[0058] wherein, the training process of the autonomous driving policy model includes:
[0059] Construct a simulation environment for autonomous driving, and formulate simple training tasks, ordinary training tasks and difficult training tasks according to the simulation environment;
[0060] Construct an autonomous driving policy model based on the DDPG algorithm;
[0061] Initialize the autonomous driving policy model and perform a first training based on the simple training task;
[0062] Perform a second training on the autonomous driving policy model after the first training based on the ordinary training task;
[0063] Perform a third training on the autonomous driving policy model after the second training based on the difficult training task.
[0064] In a third aspect, the present invention provides a device for generating an autonomous driving strategy, including a processor and a storage medium;
[0065] The storage medium is used for storing instructions;
[0066] The processor is configured to operate according to the instructions to execute the steps of the above method.
[0067] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0068] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0069] A method, device, and storage medium for generating an autonomous driving strategy provided by the present invention introduce an experience replay classification pool mechanism during the application of the DDPG algorithm. In view of the idea that in the human learning process, successful and failed experiences can be distinguished according to the learning effect for classification learning, by classifying experiences during the reinforcement learning process and extracting the leading vehicle strategy experiences and putting them into the experience classification pool for learning as well, it is possible to more specifically and reasonably learn the training cases during the autonomous driving training process, thereby improving the learning efficiency and training effect. During the reinforcement learning process, strategy reuse is added to the agent modeling process, and the agent is trained from easy to difficult, solving the problems of slow model convergence speed and poor training effect for difficult tasks in the reinforcement learning process. Compared with the original multi-agent reinforcement learning method, the method of the present invention can greatly accelerate the model convergence speed and improve the training effect for difficult tasks, and this method has strong generality and practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is a flowchart of a method for generating an autonomous driving strategy provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.
[0072] Embodiment 1:
[0073] As Figure 1 shown, the embodiment of the present invention provides a method for generating an autonomous driving strategy, including:
[0074] 1. When obtaining an autonomous driving task, initialize the vehicle's own state;
[0075] 2. Generate an autonomous driving strategy based on the trained autonomous driving strategy model according to the own state;
[0076] 3. Execute the automatic driving task according to the automatic driving strategy;
[0077] Among them, the training process of the automatic driving strategy model includes:
[0078] S1. Build a simulation environment for automatic driving, and formulate simple training tasks, ordinary training tasks, and difficult training tasks according to the simulation environment;
[0079] Specifically, the simple simulation task is the driving of a single vehicle in clear weather, the ordinary simulation task is the driving of a single vehicle in foggy weather, and the difficult simulation task is the driving of multiple vehicles in foggy weather.
[0080] S2. Build an automatic driving strategy model based on the DDPG algorithm;
[0081] S3. Initialize the automatic driving strategy model and perform one training based on the simple training task;
[0082] S3.1. Initialize the automatic driving strategy model, including initializing the real action network, target action network, real evaluation network, and target evaluation network of the automatic driving strategy model;
[0083] S3.2. Obtain a model training sample set by executing a simple training task based on the automatic driving strategy model in the simulation environment of automatic driving;
[0084] Among them, obtaining the model training sample set includes:
[0085] (1) Obtain the current state S of the vehicle under the simple training task, and bring it into the initialized automatic driving strategy model π to obtain the current action strategy A of the vehicle:
[0086] A = π(φ(S)) + N
[0087] Among them, φ(S) is the feature vector of state S, π(·) is the action strategy generated by the real action network, and N is the random noise function;
[0088] Among them, the state includes the vehicle coordinates, speed, acceleration, and the vehicle surrounding environment information. The vehicle surrounding environment information includes obstacles within a preset range around the vehicle coordinates, and the obstacles include the vehicle in front, the fence, and the traffic light.
[0089] (2) Drive the vehicle to automatically drive in the simulation environment according to the current action strategy A and obtain the next state S of the vehicle * , and obtain the action score R and the termination state E according to the next state S of the vehicle * ;
[0090] Among them, according to the next state S of the vehicle* Obtaining the action score R and the termination state E includes:
[0091] ① Obtaining the driving trajectory and driving distance of the vehicle according to the vehicle coordinates in the current state S and the next state S * ;
[0092] ② Obtaining the training base score according to the driving distance and the preset unit distance score;
[0093] ③ Judging whether the vehicle collides with an obstacle according to the driving trajectory and the obstacle coordinates. If a collision occurs, deduct the preset collision score from the training base score and set the termination state E to terminated;
[0094] ④ Judging whether the vehicle passes through the traffic light according to the driving trajectory and the traffic light coordinates. If it passes, calculate the time when the vehicle arrives at the traffic light according to the speed and acceleration of the current state, and judge the state of the traffic light according to the arrival time at the traffic light. If the traffic light is red, deduct the preset running-light score from the training base score;
[0095] ⑤ Obtaining the deducted training base score as the training score R.
[0096] (3) Saving the feature vector φ(S) of the current state of the vehicle, the feature vector φ(S * ) of the next state, the current action policy A, the training score R and the termination state E into an experience replay array, denoted as {φ(S), φ(S * ), A, R, E};
[0097] (4) Judging whether there is a vehicle in front according to the next state S * of the vehicle. If there is, obtain the experience replay array of the vehicle in front and store it in the pre-constructed experience replay set D M of the vehicle in front;
[0098] (5) Judging whether the vehicle terminates according to the termination state. If not, store the experience replay array in the pre-constructed successful experience replay set D S , and bring the next state of the vehicle into the above steps to execute in a loop;
[0099] (6) If so, store the experience replay array in the pre-constructed failed experience replay set D F , and output the failed experience replay set D F , the successful experience replay set D S and the experience replay set D M of the vehicle in front;
[0100] (7) From the successful experience replay set D S , the failed experience replay set D F and the experience replay set D MRandomly select T experience replay arrays from it and generate a model training sample set D T 。
[0101] S3.3. Train and update the autonomous driving policy model through the model training sample set, and replace the initialized autonomous driving policy model with the updated autonomous driving policy model and bring it into the above steps for iteration;
[0102] Specifically, the present invention introduces an experience replay classification pool mechanism. The idea of classifying and learning successful experiences and failed experiences according to the learning effect during the learning process. Through experience classification during the reinforcement learning process and extracting the leading vehicle policy experience and also putting it into the experience classification pool for learning, it is possible to more specifically and reasonably learn the learning cases during the autonomous driving experiment process, improving the learning efficiency and training effect. Training and updating the autonomous driving policy model through the model training sample set includes:
[0103] (1). According to the model training sample set D T Calculate the current target reward value y t :
[0104]
[0105] where y t is the target reward value of the t-th experience replay array, t = 1, 2, 3... T, R t is the training score of the t-th experience replay array, γ is the discount factor, π′(·) is the action policy generated by the target action network, ω′ is the weight parameter of the target evaluation network, Q′(·) is the evaluation value generated by the target evaluation network, is the next state of the current state S t of the t-th experience replay array;
[0106] (2) Based on the target reward value y t Construct the first loss function, and update the weight parameter ω of the real evaluation network through the gradient backpropagation of the neural network; The first loss function is:
[0107]
[0108] where ω is the weight parameter of the real evaluation network, Q(·) is the evaluation value generated by the real target evaluation network; A t is the current action policy of the t-th experience replay array;
[0109] (3) Construct the second loss function, and update the weight parameter θ of the real action network through the gradient backpropagation of the neural network; The second loss function is:
[0110]
[0111] (4) If t%C = 0, then update the weight parameters ω' and θ' of the target action network and the target evaluation network according to the weight parameters ω and θ of the real action network and the real evaluation network:
[0112] ω' ← τω + (1 - τ)ω'
[0113] θ' ← τθ + (1 - τ)θ'
[0114] Where C is the update frequency of the weight parameters of the target action network and the target evaluation network, and τ is the update coefficient;
[0115] (5) Based on the weight parameters ω, θ, ω', and θ of the updated real action network, real evaluation network, target action network, and target evaluation network ′ Update the autonomous driving policy model.
[0116] S3.4 If the preset maximum number of iterations is reached, one training is completed.
[0117] S4 Secondarily train the autonomous driving policy model after one training based on ordinary training tasks; the specific process is the same as S3, only replacing the initialized autonomous driving policy model with the autonomous driving policy model after one training, so as to perform secondary training.
[0118] S5 Tertiary train the autonomous driving policy model after secondary training based on difficult training tasks; the specific process is the same as S3, only replacing the initialized autonomous driving policy model with the autonomous driving policy model after secondary training, so as to perform tertiary training.
[0119] Embodiment 2:
[0120] The embodiment of the present invention provides an autonomous driving policy generation device, and the device includes:
[0121] An initialization module, configured to initialize the vehicle's own state when obtaining an autonomous driving task;
[0122] A policy generation module, configured to generate an autonomous driving policy based on the trained autonomous driving policy model according to the own state;
[0123] An autonomous driving module, configured to execute an autonomous driving task according to the autonomous driving policy;
[0124] Wherein, the training process of the autonomous driving policy model includes:
[0125] Construct a simulation environment for autonomous driving, and formulate simple training tasks, ordinary training tasks, and difficult training tasks according to the simulation environment;
[0126] Build an autonomous driving strategy model based on the DDPG algorithm;
[0127] Initialize the autonomous driving strategy model and conduct a single training based on a simple training task;
[0128] Conduct a second training on the autonomous driving strategy model after the single training based on a normal training task;
[0129] Conduct a third training on the autonomous driving strategy model after the second training based on a difficult training task.
[0130] Embodiment III:
[0131] Based on Embodiment I, an embodiment of the present invention provides a strategy generation device for autonomous driving, including a processor and a storage medium;
[0132] The storage medium is used to store instructions;
[0133] The processor is used to operate according to the instructions to execute the steps of the above method.
[0134] Embodiment IV:
[0135] Based on Embodiment I, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0136] A strategy generation method, device and storage medium for autonomous driving provided by the present invention:
[0137] 1. Introduce an experience replay classification pool mechanism. In view of the idea that in the human learning process, successful and failed experiences can be distinguished according to the learning effect for classification learning, by classifying experiences during the reinforcement learning process, the learning cases in the autonomous driving experiment process can be learned more pertinently, sufficiently and reasonably, improving the learning efficiency and training effect.
[0138] 2. During the strategy learning process, learn the leading vehicle strategy through an experience extraction mechanism, and put the leading vehicle strategy experience into the experience classification pool, which can imitate and learn the braking and accelerating moments of the leading vehicle during driving, avoiding the accident rate during autonomous driving.
[0139] 3. During the reinforcement learning process, combine the strategy reuse mechanism with the agent modeling process, and train the agent from easy to difficult, solving the problems of slow model convergence speed and poor training effect for difficult tasks in the reinforcement learning process.
[0140] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0141] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0144] The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A method for generating an autonomous driving strategy, characterized in that, comprising: When obtaining an autonomous driving task, initialize the vehicle's own state; Generate an autonomous driving strategy based on the trained autonomous driving strategy model according to the own state; Execute the autonomous driving task according to the autonomous driving strategy; Wherein, the training process of the autonomous driving strategy model includes: Construct a simulation environment for autonomous driving, and formulate simple training tasks, ordinary training tasks, and difficult training tasks according to the simulation environment; Construct an autonomous driving strategy model based on the DDPG algorithm; Initialize the autonomous driving strategy model and perform one-time training based on simple training tasks; Perform secondary training on the autonomous driving strategy model after one-time training based on ordinary training tasks; Perform tertiary training on the autonomous driving strategy model after secondary training based on difficult training tasks; Wherein, the initializing and performing one-time training on the autonomous driving strategy model based on simple training tasks includes: Initialize the autonomous driving strategy model, including initializing the real action network, target action network, real evaluation network, and target evaluation network of the autonomous driving strategy model; Obtain a model training sample set by executing simple training tasks based on the autonomous driving strategy model in the simulation environment of autonomous driving; Train and update the autonomous driving strategy model through the model training sample set, and replace the initialized autonomous driving strategy model with the updated autonomous driving strategy model and bring it into the above steps for iteration; If the preset maximum number of iterations is reached, one-time training is completed; Wherein, the obtaining of the model training sample set includes: Obtain the current vehicle state under a simple training task and input it into the initialized autonomous driving policy model to obtain the current action policy of the vehicle : ; Among them, is the eigenvector of the state , is the action policy generated by the realistic action network, is the random noise function; According to the current action strategy Drive the vehicle to autopilot in the simulation environment and obtain the next state of the vehicle , and according to the next state of the vehicle Obtain the training score and the termination state ; The feature vector of the current state of the vehicle , the feature vector of the next state , the current action policy , the training score and the terminal state are saved as an experience replay array, denoted as ; According to the next state of the vehicle Determine whether there is a vehicle in front. If so, obtain the experience replay array of the vehicle in front and store it in the pre-constructed experience replay set of the vehicle in front ; Determine whether the vehicle terminates according to the termination state. If not, store the experience replay array in the pre-constructed successful experience replay set and bring the next state of the vehicle into the above steps to execute in a loop; If so, store the experience replay array in the pre-constructed failed experience replay set , and output the failed experience replay set , the successful experience replay set and the leading vehicle experience replay set ; From the successful experience replay set , the failed experience replay set and the leading vehicle experience replay set , randomly extract experience replay arrays and generate a model training sample set .
2. A method for generating an autonomous driving strategy according to claim 1, characterized in that, The simple training task is a single vehicle driving in clear weather, the ordinary training task is a single vehicle driving in foggy weather, and the difficult training task is multiple vehicles driving in foggy weather.
3. A method for generating an autonomous driving strategy according to claim 1, characterized in that, The state of the vehicle includes vehicle coordinates, speed, acceleration, and vehicle surrounding environment information. The vehicle surrounding environment information includes obstacles within a preset range around the vehicle coordinates. The obstacles include the vehicle in front, the fence, and the traffic signal.
4. A method for generating an autonomous driving strategy according to claim 3, characterized in that, According to the next state of the vehicle Obtain a training score And a termination state Including: Based on the current state and the next state obtain the driving trajectory and driving distance of the vehicle according to the vehicle coordinates in them; Obtain a training base score according to the driving distance and the preset unit distance score; Determine whether the vehicle collides with an obstacle based on the driving trajectory and the obstacle coordinates. If a collision occurs, deduct a preset collision score from the training base score and set the state to terminated. Set it to terminated. Judge whether the vehicle passes the traffic signal according to the driving trajectory and the traffic signal coordinates. If it passes, calculate the time when the vehicle arrives at the traffic signal according to the speed and acceleration of the current state, and judge the state of the traffic signal according to the arrival time of the traffic signal. If the traffic signal is red, deduct the preset red light violation score from the training base score; Obtain the training base score after deduction as the training score .
5. A method for generating an autonomous driving strategy according to claim 1, characterized in that, The training and updating of the autonomous driving strategy model through the model training sample set includes: According to the model training sample set Calculate the current target reward value : ; Among them, is the target reward value of the th experience replay array, , is the training score of the th experience replay array, is the discount factor, is the action policy generated by the target action network, are the weight parameters of the target evaluation network, is the evaluation value generated by the target evaluation network, is the current state of the th experience replay array 's next state; Based on the target reward value Construct a first loss function and update the weight parameters of the reality evaluation network through the gradient backpropagation of the neural network ; The first loss function is as follows: ; Among them, is the weight parameter of the reality evaluation network, is the evaluation value generated by the reality target evaluation network; is the current action strategy of the Construct a second loss function, and update the weight parameters of the real action network through gradient backpropagation of the neural network ; The second loss function is as follows: ; If , then update the weight parameters of the target action network and the target evaluation network according to the weight parameters of the real action network and the real evaluation network and : and : ; ; Among them, is the update frequency of the weight parameters of the target action network and the target evaluation network, is the update coefficient; Based on the weight parameters of the updated reality action network, reality evaluation network, target action network, and target evaluation network , , and update the autonomous driving policy model.
6. An apparatus for generating an autonomous driving strategy, characterized in that, The apparatus includes: An initialization module, configured to initialize the vehicle's own state when obtaining an autonomous driving task; A strategy generation module for generating an autonomous driving strategy based on the trained autonomous driving strategy model according to its own state; An autonomous driving module for performing an autonomous driving task according to the autonomous driving strategy; Wherein, the training process of the autonomous driving strategy model includes: Constructing a simulation environment for autonomous driving, and formulating simple training tasks, ordinary training tasks, and difficult training tasks according to the simulation environment; Constructing an autonomous driving strategy model based on the DDPG algorithm; Initializing the autonomous driving strategy model and performing a first training based on the simple training task; Performing a second training on the autonomous driving strategy model after the first training based on the ordinary training task; Performing a third training on the autonomous driving strategy model after the second training based on the difficult training task; Wherein, the initializing the autonomous driving strategy model and performing a first training based on the simple training task includes: Initializing the autonomous driving strategy model, including initializing the real action network, target action network, real evaluation network, and target evaluation network of the autonomous driving strategy model; Obtaining a model training sample set by executing a simple training task based on the autonomous driving strategy model in the simulation environment of autonomous driving; Training and updating the autonomous driving strategy model through the model training sample set, and substituting the updated autonomous driving strategy model for the initialized autonomous driving strategy model into the above steps for iteration; If the preset maximum number of iterations is reached, the first training is completed; Wherein, the obtaining the model training sample set includes: Obtain the current state of the vehicle under a simple training task and input it into the initialized autonomous driving policy model to obtain the current action policy of the vehicle : ; Among them, is the feature vector of the state , is the action policy generated by the realistic action network is the random noise function; According to the current action strategy Drive the vehicle to drive automatically in the simulation environment and obtain the next state of the vehicle , and according to the next state of the vehicle Obtain the training score And the termination state ; The feature vector of the current state of the vehicle , the feature vector of the next state , the current action policy , the training score and the termination state are saved as an experience replay array, denoted as ; According to the next state of the vehicle Determine whether there is a vehicle in front. If so, obtain the experience replay array of the vehicle in front and store it in the pre-constructed experience replay set of the vehicle in front ; Determine whether the vehicle terminates according to the termination status. If not, store the experience replay array in the pre-constructed successful experience replay set and bring the next state of the vehicle into the above steps for loop execution; If so, store the experience replay array in a pre-constructed failed experience replay set , and output the failed experience replay set , the successful experience replay set and the leading vehicle experience replay set ; From the successful experience replay set , the failed experience replay set and the leading vehicle experience replay set , randomly extract experience replay arrays and generate a model training sample set .
7. A device for generating an autonomous driving strategy, Characterized in that, It includes a processor and a storage medium; The storage medium is used for storing instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1-5.
8. A computer-readable storage medium, on which a computer program is stored, Characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Automobile traffic safety simulation driving education training system
CN102184659A
Anthropomorphic automatic driving car-following model based on deep reinforcement learning
CN109733415A
Automatic driving method and device
CN111984018A
Multi-legged robot motion control method based on depth deterministic strategy gradient
CN113031528A