Lane-changing Decision-making System and Method for Autonomous Driving Vehicles Based on Deep Reinforcement Learning

Through the decision-making system for lane change of autonomous vehicles based on deep reinforcement learning, lane change strategies are optimized and trained, the problems of insufficient data volume and limited scenarios during lane change are solved, and the accuracy and safety of lane change are improved.

CN114802248BActive Publication Date: 2025-06-20INTELLIGENT CONNECTED TECH OF CAERI CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210443895.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-25
Publication Date
2025-06-20
Estimated Expiration
2042-04-25

AI Technical Summary

Technical Problem

There are problems such as insufficient data volume and limited road change scenarios during the lane change process, resulting in low accuracy and safety of lane change.

Method used

The lane change decision system of autonomous driving vehicles based on deep reinforcement learning is adopted. Through the processor module, data acquisition module, data analysis module and lane change strategy module, the Actor-Critic algorithm and Markov decision-making process are used to optimize and train the lane change strategy to generate a more accurate lane change strategy.

Benefits of technology

It improves the accuracy and safety of autonomous vehicles during lane change, reduces road congestion and collision accidents, and can make correct and reasonable responses to new traffic scenarios under limited data conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114802248B_ABST
    Figure CN114802248B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of autonomous driving control technology, and discloses an autonomous driving vehicle lane-changing decision-making system and method based on deep reinforcement learning, including a processor module, as well as a data acquisition module, a data analysis module, and a lane-changing strategy module respectively connected to the processor module. By analyzing the data information of the target vehicle collected, as well as the operation data of the interfering vehicles near the target vehicle, a safe autonomous lane-changing strategy for the current target vehicle is obtained, so as to achieve a fast and safe lane change. The present invention has the beneficial effects of improving the accuracy of the autonomous lane-changing strategy of autonomous driving vehicles, ensuring the safety of vehicles during the lane-changing process, reducing road congestion and the occurrence of collision accidents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving control, and particularly to a lane-changing decision-making system and method for an autonomous driving vehicle based on deep reinforcement learning. Background Art

[0002] In recent years, autonomous driving has received particular attention worldwide and is regarded as an important technology to alleviate traffic congestion, reduce traffic accidents and environmental pollution. At present, some autonomous driving vehicles have carried out large-scale road tests, such as Google's autonomous driving and Apple's autonomous driving. According to research, in current traffic accidents, more than 30% of road accidents are caused by unreasonable lane-changing behaviors. Therefore, the research on lane-changing assistance technology in intelligent assisted driving technology is particularly important. At present, the mainstream rule-based algorithms are faced with the problem that the insufficient data volume leads to the model being unable to fully handle the infinite scenarios of autonomous driving vehicles during the lane-changing process, resulting in lane-changing failure or affecting the safety during the lane-changing process.

[0003] In the prior art, there is a lane-changing decision-making control method for an autonomous driving vehicle based on hierarchical reinforcement learning, which belongs to the technical field of autonomous driving control. It solves the problems of poor safety / low efficiency existing in the prior autonomous driving process. The present invention uses the speed in the actual driving scenario of the autonomous driving vehicle and the relative position and relative speed information of the vehicles in the surrounding environment to establish a decision neural network with 3 hidden layers, and uses a lane-changing safety reward function to train and fit the Q valuation function of the decision neural network to obtain the action with the maximum Q valuation; uses the speed in the actual driving scenario of the autonomous driving vehicle and the relative position information of the surrounding environment vehicles and the reward function corresponding to the following or lane-changing action to establish a deep Q learning acceleration decision model to obtain lane-changing or following acceleration information. When lane-changing, a 5th-degree polynomial curve is used to generate a reference lane-changing trajectory. The present invention is applicable to autonomous driving lane-changing decision-making and control.

[0004] Although this solution can simulate and train for lane-changing and determine the reward function for the scenario and the surrounding environment to ensure the training results of autonomous driving lane-changing, it still has the problems of insufficient data volume and limited lane-changing scenarios for specific lane-changing scenarios, ultimately resulting in too low accuracy and safety of its lane-changing. Summary of the Invention

[0005] The present invention aims to provide a lane-changing decision-making system and method for an autonomous driving vehicle based on deep reinforcement learning to improve the accuracy of the lane-changing strategy of the autonomous driving vehicle and ensure lane-changing safety.

[0006] To achieve the above object, the present invention adopts the following technical solutions: A lane-changing decision-making system for an autonomous driving vehicle based on deep reinforcement learning, comprising a processor module, and a data acquisition module, a data analysis module, and a lane-changing strategy module respectively connected to the processor module;

[0007] The data acquisition module is used to collect the data information of the target vehicle and the running data of the interfering vehicles near the target vehicle, and then form a first data set and send the first data set to the data analysis module;

[0008] The data analysis module is used to analyze and process the first data set, and obtain the lane-changing scenario and lane-changing data of the autonomous driving vehicle;

[0009] The lane-changing strategy module is used to generate a first lane-changing strategy according to the obtained lane-changing scenario and lane-changing data, and send the first lane-changing strategy to the processor module;

[0010] The processor module includes a data storage unit and a lane-changing execution unit. The data storage unit is used to store the first lane-changing strategy; the lane-changing execution unit is used to obtain a rule-based lane-changing trajectory execution model according to the first lane-changing strategy and control the autonomous driving vehicle to change lanes.

[0011] The principle and advantages of this solution are: In practical applications, based on the rule-based lane-changing model, the deep reinforcement learning method is used to train and attempt the lane-changing model, and the Actor-Critic algorithm is used to continuously optimize the lane-changing strategy of the autonomous driving vehicle, so that the autonomous driving vehicle can accurately handle the problem of infinite scenarios during lane-changing, thereby improving the accuracy of the automatic lane-changing strategy of the autonomous driving vehicle, ensuring the safety of the vehicle during lane-changing, and reducing road congestion and collision accidents. Compared with the prior art, the advantage of the present invention is that the established deep learning model can test the lane-changing strategy of autonomous driving technology more comprehensively and accurately, guide the autonomous driving vehicle to complete lane-changing quickly and safely, and the obtained lane-changing trajectory can be applied to the lane-changing scenario of the autonomous driving vehicle, so that the vehicle can make correct and reasonable responses to new traffic scenarios under limited data volume conditions, ensuring driving safety.

[0012] Preferably, as an improvement, the first data set includes the surrounding vehicle information at the current moment, the surrounding road information at the current moment, the surrounding road information at the next moment, the surrounding vehicle information at the next moment, and the vehicle information of the vehicle itself.

[0013] Beneficial effects: By collecting the surrounding vehicle information and road information, it can accurately provide the data for reference in lane-changing, so as to make a safety judgment for lane-changing, avoid collision safety accidents between the target vehicle and surrounding interfering vehicles during lane-changing, and at the same time can greatly improve the accuracy of the lane-changing strategy of this lane-changing model.

[0014] Preferably, as an improvement, the analysis and processing of the first data set is to use a preset analysis algorithm to perform infinite scenario exploration and analysis on the finite first data set, and perform deep reinforcement learning on the analysis process before obtaining the corresponding lane-changing scenario.

[0015] Beneficial effects: By analyzing and processing the collected data through this process, the defect of insufficient current data volume can be effectively overcome, and infinite scenario analysis of the lane-changing model and lane-changing strategy can be performed on the limited data volume, which can greatly improve the diversity of lane-changing scenarios and the accuracy of lane-changing strategies, providing a reliable guarantee for the automatic and safe lane-changing behavior of autonomous vehicles, thereby ensuring driving safety.

[0016] Preferably, as an improvement, the preset analysis algorithm is the Actor-Critic algorithm; the deep reinforcement learning is to use the Markov decision process to describe the analysis process, forming a six-tuple M = (S, A, P, r, ρ, γ), where S is the state space, and the state space is the set of all states; A is the action space, and the action space is the set of all actions; P is the state transition probability; r is the reward function of the state transition process; γ is the discount factor in the state transition process.

[0017] Beneficial effects: Through the interaction between the data and the environment strengthened by the Actor-Critic algorithm, and the learning does not simply obtain the optimal strategy for single-step decision-making, but pursues the long-term cumulative reward obtained from the interaction with the environment, making the training result of the lane-changing strategy for autonomous vehicles more accurate and ensuring the safety of lane-changing.

[0018] Preferably, as an improvement, the reward function is where v is the real-time speed of the vehicle, v min is the minimum speed adopted during the vehicle training process, v max is the maximum speed adopted during the vehicle training process, a is the speed reward value for the lane-changing process, b is the collision penalty value for the vehicle collision, and collision is the feedback result of the simulation environment for the vehicle collision.

[0019] Beneficial effects: During the training process of the lane-changing model, the reward function is used to accumulate rewards for the training process, and the parameter data of the lane-changing model is corrected, thereby improving the accuracy of the lane-changing model.

[0020] Preferably, as an improvement, the data analysis module uses the deep reinforcement learning method to train and attempt the lane-changing model based on the rule-based lane-changing model, and finally verifies the model.

[0021] Beneficial effects: For the lane-changing model obtained after multiple reinforcement learnings, to ensure accuracy, the method of deep reinforcement learning is used for training and experimentation, thereby completing the correction of the lane-changing model, ensuring the correctness of the lane-changing model parameters, and providing a more accurate lane-changing strategy service for autonomous driving vehicles.

[0022] Preferably, as an improvement, when generating a lane-changing strategy, the lane-changing strategy module uses a rule-based trajectory planning algorithm to assist in the calculation. The expression of the rule-based trajectory planning algorithm is where θ i is the heading angle at the starting point of the planning step length, is the lateral coordinate of the end point, x n is the longitudinal position of vehicle n, y n is the lateral position of vehicle n.

[0023] Beneficial effects: By planning the lane-changing trajectory in this way, the safety of autonomous driving vehicles during lane-changing is ensured, collisions with other vehicles are avoided, and the lane-changing time of the vehicle can be effectively reduced, the traffic efficiency of road vehicles can be improved, and road congestion can be effectively alleviated.

[0024] Preferably, as an improvement, when using the Markov decision process to describe the analysis process, the state value function is defined as follows. where a t , r t , s t+1 , a t+1 , r t+1 ,...~π indicates that the trajectory comes from the interaction between the policy π and the environment.

[0025] Beneficial effects: By continuously updating the road environment data, new lane-changing decision data is provided, ensuring the continuous update of new data, thereby improving the real-time performance and accuracy of the lane-changing model, and ensuring the lane-changing safety of autonomous driving vehicles.

[0026] The present invention also provides a lane-changing decision method for autonomous driving vehicles based on deep reinforcement learning, including the following steps:

[0027] Step S1, collecting the data information of the target vehicle and the operation data of the interfering vehicles near the target vehicle;

[0028] Step S2, using the Actor-Critic algorithm to analyze and process the collected data, and combining the Markov decision process to describe the reinforcement learning problem to obtain the lane-changing scenario and lane-changing data of the autonomous driving vehicle;

[0029] Step S3, using the rule-based trajectory planning algorithm to calculate the lane-changing trajectory and using the reward function to correct the test process to finally obtain the lane-changing strategy;

[0030] Step S4: Verify the output results of this model using the autonomous driving simulation environment highway_env and the dynamics-based simulation software CarSim.

[0031] Beneficial effects: Using this method to implement the testing of the automatic lane-changing model for autonomous vehicles and the generation of lane-changing strategies can ensure the lane-changing safety of vehicles, improve the lane-changing efficiency, reduce road congestion, and also greatly improve road traffic safety.

[0032] Preferably, as an improvement, verifying the output results of this model means using the highway scenario in the autonomous driving simulation environment to test the vehicle lane-changing strategy, and using two statistics, the mean absolute error and the mean absolute relative error, to perform error statistics on the model.

[0033] Beneficial effects: By verifying the output results of the model, it can ensure the decision-making accuracy of the model for the lane-changing strategy to the greatest extent, thus avoiding collisions between lane-changing vehicles and surrounding vehicles, improving the safety and efficiency of lane-changing, and on the other hand, also improving the traffic efficiency of road traffic and alleviating the congestion situation. Description of the Drawings

[0034] Figure 1 It is a system schematic diagram of Embodiment 1 of the lane-changing decision-making system for autonomous vehicles based on deep reinforcement learning of the present invention.

[0035] Figure 2 It is a schematic diagram of the LSTM neural network of Embodiment 1 of the lane-changing decision-making system for autonomous vehicles based on deep reinforcement learning of the present invention.

[0036] Figure 3 It is a schematic diagram of the composition of the LSTM neural network of Embodiment 1 of the lane-changing decision-making system for autonomous vehicles based on deep reinforcement learning of the present invention.

[0037] Figure 4 It is a schematic flowchart of Embodiment 1 of the lane-changing decision-making method for autonomous vehicles based on deep reinforcement learning of the present invention.

[0038] Figure 5 It is a schematic diagram of the profit change of Embodiment 1 of the lane-changing decision-making method for autonomous vehicles based on deep reinforcement learning of the present invention. Detailed Embodiments

[0039] The following is a further detailed description through specific embodiments:

[0040] The marks in the attached drawings of the specification include: processor module 1, data acquisition module 2, data analysis module 3, lane-changing strategy module 4, data storage unit 5, lane-changing execution unit 6.

[0041] Embodiment 1:

[0042] This embodiment is basically as shown in the appendix Figure 1 : A lane-changing decision-making system for an autonomous driving vehicle based on deep reinforcement learning, including a processor module 1, and a data acquisition module 2, a data analysis module 3, and a lane-changing strategy module 4 that are respectively connected to the processor module 1;

[0043] The data acquisition module 2 is used to collect the data information of the target vehicle and the running data of the interfering vehicles near the target vehicle, and then form a first data set and send the first data set to the data analysis module 3;

[0044] The data analysis module 3 is used to analyze and process the first data set, and obtain the lane-changing scenario and lane-changing data of the autonomous driving vehicle;

[0045] The lane-changing strategy module 4 is used to generate a first lane-changing strategy according to the obtained lane-changing scenario and lane-changing data, and send the first lane-changing strategy to the processor module 1;

[0046] The processor module 1 includes a data storage unit 5 and a lane-changing execution unit 6. The data storage unit 5 is used to store the first lane-changing strategy; the lane-changing execution unit 6 is used to obtain a rule-based lane-changing trajectory execution model according to the first lane-changing strategy and control the autonomous driving vehicle to change lanes.

[0047] The data collected by the data acquisition module 2 includes the surrounding vehicle information at the current moment, the surrounding road information at the current moment, the surrounding road information at the next moment, the surrounding vehicle information at the next moment, and the vehicle information of the vehicle itself;

[0048] As shown in the appendix Figure 2 : The first data set is processed using the Actor-Critic algorithm, and the Markov decision process is used to describe the reinforcement learning problem, forming a six-dimensional tuple M=(S, A, P, r, ρ, γ), where S is the state space, that is, the set of all states; A is the action space, that is, the set of all actions; P is the state transition probability; r is the reward function of the state transition process; γ is the discount factor in the state transition process. Under the Markov decision process M and the policy π, the state value function is defined as follows,

[0049]

[0050] where, a t ,r t ,s t+1 ,a t+1 ,r t+1,...~π represents that the trajectory comes from the interaction between the policy π and the environment. The above formula represents the expected cumulative reward obtained by the agent through interacting with the environment using the policy π starting from the state st = s. Similarly, the state-action value function can also be defined:

[0051]

[0052] It represents the expected cumulative reward obtained by the agent through interacting with the policy π after executing the action at starting from the state st, and the state value function and the state-action value function can be converted to each other. When the policy π is a probabilistic policy,

[0053]

[0054]

[0055] For any probabilistic policy π, there is the following Bellman expectation equation,

[0056]

[0057]

[0058] When the agent adopts the policy π, after executing the action at from the state st and transferring to the state st+1 and obtaining the reward rt, the following dynamic programming update is directly performed,

[0059] V π (S t ) = V π (S t ) + α(r t + γV(S t+1 ) - V(S t ))

[0060] Let r t + γV(s t+1 ) - V(s t ) = TD - error

[0061] After obtaining, the neural network will perform backpropagation. The Actor is based on the policy gradient. The policy is parameterized as a neural network, denoted by θ. The direction of θ iteration is to maximize the expectation of the episodic reward, and the objective function is expressed as:

[0062]

[0063] Among them, τ represents a sampling period, and π θ (τ) represents the probability of the sequence appearance. Taking the gradient of J(θ) gives:

[0064]

[0065] Then:

[0066]

[0067] Finally, update the neural network parameters:

[0068]

[0069] As shown in the appendix Figure 3 The LSTM neural network mainly consists of input layer, hidden layer, and output layer neurons. The hidden layer neurons mainly have three gate structures and one state: forget gate, input gate, output gate, and cell state. For the later training and correction of the lane-changing strategy, first, when new data is input into the long short-term memory network, it is necessary to determine which old data needs to be discarded from the cell state, and this part is determined by the forget gate, which is a sigmoid function layer

[0070]

[0071] In the formula, W f is the weight matrix of the forget gate, h t-1 is the cell state at time t-1, x t is the environmental input data, b f is the bias term of the forget gate

[0072] Then, passing through a sigmoid function layer, that is, the input gate will determine which values need to be updated, and then a tanh function layer will create a vector as the candidate value to be added to the cell state:

[0073]

[0074]

[0075] In the formula, b i is the bias term of the input gate is the data matrix to be updated, W c is the weight matrix of the data to be updated

[0076] Then, update the cell state of the previous moment. First, remove the information determined by the forget gate from the cell state, and then add the candidate value calculated by the input gate at the ratio of the proportion of each state value to be updated

[0077]

[0078] Finally, determine the part to be output. The output is appropriately processed based on the cell state, that is, through a sigmoid function layer to determine which parts need to be updated, and then processed through a tanh function and multiplied by the output of the sigmoid layer in the forget gate to determine the output:

[0079] O t = σ(W o [h t-1 , x t + b o )

[0080] Where W o is the weight matrix of the output gate, and b o is the bias term of the output gate.

[0081] The cell state before improvement is:

[0082] s t = tanh(W g [h t-1 , x t + b g ) · σ(W i [h t-1 , x t + b i ) + s t-1 · σ(W f [h t-1 , x t b f ))

[0083] The output is:

[0084] h t = tanh(s t ) · σ(W o [h t-1 , x t + b o )

[0085] Then the error between the value of the p-th output and the true value is:

[0086]

[0087] Select a cubic polynomial with a relatively simple form as the rule-based trajectory planning algorithm, and its expression is as follows:

[0088]

[0089] Where θ i is the heading angle at the starting point of the planning step length, is the end horizontal coordinate, xn is the longitudinal position of vehicle n, and yn is the horizontal position of vehicle n.

[0090] As shown in the appendix Figure 4 The present invention also provides a lane-changing decision-making method for an autonomous driving vehicle based on deep reinforcement learning applied to the above system, including the following steps:

[0091] Step S1, collect the data information of the target vehicle and the running data of the interfering vehicles near the target vehicle;

[0092] Step S2, use the Actor-Critic algorithm to analyze and process the collected data, and combine the Markov decision process to describe the reinforcement learning problem, to obtain the lane-changing scenario and lane-changing data of the autonomous driving vehicle;

[0093] Step S3, use the rule-based trajectory planning algorithm to calculate the lane-changing trajectory and use the reward function to correct the test process, and finally obtain the lane-changing strategy;

[0094] Step S4, obtain a rule-based lane-changing trajectory execution model according to the lane-changing strategy, and use the autonomous driving simulation environment highway_env and the dynamics-based simulation software CarSim to verify the output results of this model.

[0095] During the reinforcement learning process, use the autonomous driving simulation environment to test the vehicle lane-changing strategy using the highway scenario, and use two statistics, the mean absolute error and the mean absolute relative error, to perform error statistics on the model.

[0096]

[0097] In the formula, N represents the number of test data samples, d r,i represents the nominal value of the i-th vehicle, and d s,i represents the predicted value of the i-th vehicle.

[0098] As shown in the appendix Figure 5 It can be seen from the trend of the change of the model revenue with the increase of the training times that the training revenue rises rapidly with the increase of the training times. When the training exceeds 2000 times, the revenue value is stable and tends to converge.

[0099] Using this system, it is possible to update the vehicle lane-changing strategy based on the actual current data, so as to provide an automatic lane-changing strategy and method for autonomous driving vehicles, thereby ensuring the smooth progress of the lane-changing process, avoiding collisions with surrounding vehicles, improving the safety and efficiency of lane-changing, further ensuring the smooth flow of road traffic, and reducing urban traffic congestion.

[0100] The specific implementation process of this embodiment is as follows:

[0101] First step: Use the data acquisition module 2 to collect the data information of the target vehicle and the running data of the interfering vehicles near the target vehicle, including the information of surrounding vehicles at the current moment, the information of surrounding roads at the current moment, the information of surrounding roads at the next moment, the information of surrounding vehicles at the next moment, and the information of the vehicle itself. Then form the first data set and send the first data set to the data analysis module 3.

[0102] Second step: After receiving the first data set, the data analysis module 3 analyzes and processes the first data set. Use the Actor-Critic algorithm to process the first data set, and use the Markov decision process to describe the reinforcement learning problem, forming a six-dimensional tuple M=(S, A, P, r, ρ, γ). Then obtain the lane-changing scenario and lane-changing data of the autonomous driving vehicle.

[0103] Third step: The lane-changing strategy module 4 generates the first lane-changing strategy according to the obtained lane-changing scenario and lane-changing data, and sends the first lane-changing strategy to the processor module 1. The data storage unit 5 of the processor module 1 receives and stores the first lane-changing strategy, and the lane-changing execution unit 6 obtains a rule-based lane-changing trajectory execution model according to the first lane-changing strategy and controls the autonomous driving vehicle to change lanes.

[0104] Fourth step: Train and test the lane-changing trajectory execution model. Finally, use the autonomous driving simulation environment highway_env and the dynamics-based simulation software CarSim to verify the output results of this model. Use the autonomous driving simulation environment to test the vehicle lane-changing strategy using highway scenarios, and use two statistics, the mean absolute error and the mean absolute relative error, to statistically analyze the errors of the model. Finally, the obtained lane-changing trajectory and speed change smoothly in the dynamics simulation and can be tracked under the condition of maintaining a small error with the target trajectory, and the vehicle driving stability is good.

[0105] In recent years, autonomous driving has received particular attention worldwide and is considered an important technology to alleviate traffic congestion, reduce traffic accidents and environmental pollution. Currently, some autonomous driving has conducted large-scale road tests, such as Google's autonomous driving and Apple's autonomous driving. According to research, in current traffic accidents, more than 30% of road accidents are caused by unreasonable lane-changing behaviors. Therefore, the research on the lane-changing assistance technology in intelligent assisted driving technology is particularly important. At present, the mainstream rule-based algorithms are faced with the problem that the insufficient data volume leads to the model being unable to fully handle the infinite scenarios of autonomous driving vehicles during lane-changing, resulting in lane-changing failures or affecting the safety during the lane-changing process.

[0106] In this solution, a deep reinforcement learning model is established, and the deficiencies of current mainstream lane-changing models are comprehensively considered. Thus, a rule-based training model is introduced, presenting a solution idea for how a neural network can adapt to and learn more driving skills with limited data. Using the data information of the target vehicle collected and the operation data of the interfering vehicles near the target vehicle for statistical analysis and processing, the Actor-Critic algorithm is used to process the first data set, and the Markov decision process is used to describe the reinforcement learning problem, forming a six-dimensional tuple M = (S, A, P, r, ρ, γ). Then, the lane-changing model is trained and tested using the deep reinforcement learning method. Finally, the output results of this model are verified using the autonomous driving simulation environment highway_env and the dynamics-based simulation software CarSim. Throughout the verification process, compared with the past, since the lane-changing scenarios and the amount of selected lane-changing strategy data in this model are larger, and the lane-changing strategy provided by this solution is more accurate, the lane-changing travel time in this solution is actually less than in the past. It is found by comparison that the travel time is reduced by more than 50%, and the collision probability is reduced by more than 75%. Furthermore, the authenticity and accuracy of the lane-changing strategy of this model are effectively ensured, guiding the autonomous driving vehicle to complete the lane change quickly and safely. The obtained lane-changing trajectory can be applied to the lane-changing scenario of the autonomous driving vehicle, enabling the vehicle to make a reasonable response to the new traffic scenario under the condition of limited data volume. Not only can a large amount of backup lane-changing strategy data be provided, but also the monitoring data of the surrounding environment during the lane change is continuously fed back to this model in real time, so as to timely interfere and follow up the lane-changing strategy during the lane change, improving the safety of the autonomous driving vehicle's lane change, effectively reducing the road traffic congestion and accident incidence rate, and ensuring the safety of the passengers and drivers.

[0107] Embodiment 2:

[0108] This embodiment is basically the same as Embodiment 1, except that: this system further includes a display module. During the process of analyzing and establishing the lane-changing model, the display module is used to display the lane-changing trajectory and vehicle operation data of the autonomous driving vehicle in real time, and the specific situation of the vehicle's lane change can be known more intuitively and accurately.

[0109] The specific implementation process of this embodiment is the same as that of Embodiment 1, except that:

[0110] In the fourth step, the lane-changing trajectory execution model is trained and tested. Finally, the output results of this model are verified using the autonomous driving simulation environment highway_env and the dynamics-based simulation software CarSim. The vehicle lane-changing strategy is tested using the highway scenario in the autonomous driving simulation environment, and the error of the model is statistically analyzed using two statistics, the mean absolute error and the mean absolute relative error. Finally, an accurate autonomous driving lane-changing strategy is obtained. During the entire training and verification process, the display module is used to display the lane-changing trajectory and vehicle operation data of the autonomous driving vehicle in real time.

[0111] Provide the function of displaying the lane-changing trajectory and lane-changing data, so that the lane-changing test of the autonomous driving vehicle is more accurate, and the operator can more intuitively understand the entire test process, which is convenient for correcting unqualified places and improving the test efficiency of the lane-changing test.

[0112] Embodiment 3:

[0113] This embodiment is basically the same as Embodiment 1, except that: the system further includes a lane-changing trajectory correction module, which is used to correct the lane-changing trajectory of the autonomous driving vehicle when it cannot safely complete the lane change according to the current lane-changing trajectory due to changes in road condition information during the lane-changing process, so that the autonomous driving vehicle can successfully complete the lane change, improve the safety guarantee of the lane change, and at the same time can greatly guarantee the safety of the driver and passengers.

[0114] The specific implementation process of this embodiment is the same as that of Embodiment 1, except that:

[0115] In the third step, the lane-changing strategy module 4 generates a first lane-changing strategy according to the obtained lane-changing scenario and lane-changing data, and sends the first lane-changing strategy to the processor module 1. The data storage unit 5 of the processor module 1 receives and stores the first lane-changing strategy, and the lane-changing execution unit 6 obtains a rule-based lane-changing trajectory execution model according to the first lane-changing strategy and controls the autonomous driving vehicle to change lanes; during the lane-changing process, if the road condition information changes due to the following unexpected situations, resulting in the inability to safely complete the lane change according to the current lane-changing trajectory, the lane-changing trajectory correction module corrects the lane-changing trajectory of the autonomous driving vehicle in real time, so that the autonomous driving vehicle can successfully complete the lane change.

[0116] Considering the irregular changes in the surrounding road environment of the autonomous driving vehicle and the emergency braking or sudden acceleration of surrounding vehicles, etc., during the process of the autonomous driving vehicle changing lanes according to the current lane-changing strategy, if the road information changes, the lane-changing trajectory correction module intervenes to correct the lane-changing trajectory in real time, so as to ensure the safe and smooth progress of the lane change, not only guarantee the safety of the lane change, but also ensure the safety of the driver and passengers, reduce the occurrence of traffic accidents, and improve the traffic situation.

[0117] The above are only embodiments of the present invention, and common general technical solutions and / or characteristics in the solution are not described in detail herein. It should be noted that for those skilled in the art, without departing from the technical solution of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent. The protection scope required by this application shall be subject to the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims.

Claims

1. A lane-changing decision-making system for autonomous driving vehicles based on deep reinforcement learning, characterized in that: It includes a processor module, as well as a data acquisition module, a data analysis module, and a lane-changing strategy module that are respectively connected to the processor module; The data acquisition module is used to collect the data information of the target vehicle and the running data of the interfering vehicles near the target vehicle, and then form a first data set and send the first data set to the data analysis module; The data analysis module is used to analyze and process the first data set, and obtain the lane-changing scenario and lane-changing data of the autonomous vehicle; The lane-changing strategy module is used to generate a first lane-changing strategy according to the obtained lane-changing scenario and lane-changing data, and send the first lane-changing strategy to the processor module; The processor module includes a data storage unit and a lane-changing execution unit. The data storage unit is used to store the first lane-changing strategy; The lane-changing execution unit is used to obtain a rule-based lane-changing trajectory execution model according to the first lane-changing strategy and control the autonomous vehicle to change lanes; The analysis and processing of the first data set is to use a preset analysis algorithm to perform infinite scenario exploration analysis on the limited first data set, and perform deep reinforcement learning on the analysis process before obtaining the corresponding lane-changing scenario; The preset analysis algorithm is the Actor-Critic algorithm; the deep reinforcement learning is to use the Markov decision process to describe the analysis process, forming a six-tuple M = (S, A, P, R, ρ, γ), where S is the state space, and the state space is the set of all states; A is the action space, and the action space is the set of all actions; P is the state transition probability; R is the reward function of the state transition process; γ is the discount factor in the state transition process; The reward function is Where v is the real-time speed of the vehicle, is the minimum speed adopted during the vehicle training process, is the maximum speed adopted during the vehicle training process, a is the speed reward value for the lane-changing process, b is the collision penalty value for vehicle collisions, and collision is the feedback result of the simulation environment for vehicle collisions.

2. The lane-changing decision-making system for autonomous driving vehicles based on deep reinforcement learning according to claim 1, characterized in that: The first data set includes the surrounding vehicle information at the current moment, the surrounding road information at the current moment, the surrounding road information at the next moment, the surrounding vehicle information at the next moment, and the vehicle information of the host vehicle.

3. The lane-changing decision-making system for autonomous driving vehicles based on deep reinforcement learning according to claim 1, characterized in that: The data analysis module uses the deep reinforcement learning method to train and attempt the lane-changing model on the basis of the rule-based lane-changing model, and finally verifies the model.

4. The lane-changing decision-making system for autonomous driving vehicles based on deep reinforcement learning according to claim 1, characterized in that: The lane-changing strategy module uses a rule-based trajectory planning algorithm to assist in the calculation when generating the lane-changing strategy. The expression of the rule-based trajectory planning algorithm is wherein is the course angle at the starting point of the planned step length, is the transverse coordinate of the end point, is the longitudinal position of vehicle n, is the transverse position of vehicle n.

5. The lane-changing decision-making system for autonomous driving vehicles based on deep reinforcement learning according to claim 1, characterized in that: When using the Markov decision process to describe the analysis process, the state value function is defined as follows: where indicates that the trajectory comes from the interaction between policy π and the environment.

6. A lane-changing decision-making method for autonomous driving vehicles based on deep reinforcement learning, characterized in that: It includes the following steps: Step S1, collect the data information of the target vehicle and the running data of the interfering vehicles near the target vehicle; Step S2, use the Actor-Critic algorithm to analyze and process the collected data, and combine the Markov decision process to describe the reinforcement learning problem, and obtain the lane-changing scenario and lane-changing data of the autonomous vehicle; Use the Markov decision process to describe the analysis process, forming a six-tuple M = (S, A, P, R, ρ, γ), where S is the state space, and the state space is the set of all states; A is the action space, and the action space is the set of all actions; P is the state transition probability; R is the reward function of the state transition process; γ is the discount factor in the state transition process; The reward function is where v is the real-time speed of the vehicle, is the minimum speed adopted during the vehicle training process, is the maximum speed adopted during the vehicle training process, a is the speed reward value for the lane-changing process, b is the collision penalty value for vehicle collisions, and collision is the feedback result of the simulation environment for vehicle collisions; Step S3: Use a rule-based trajectory planning algorithm to calculate the lane-changing trajectory and use the reward function to correct the testing process, and finally obtain the lane-changing strategy; Step S4: Obtain a rule-based lane-changing trajectory execution model according to the lane-changing strategy, and use the autonomous driving simulation environment highway_env and the dynamics-based simulation software CarSim to verify the output results of this model.

7. The lane-changing decision-making method for autonomous driving vehicles based on deep reinforcement learning according to claim 6, characterized in that: The verification of the output results of this model is to use the autonomous driving simulation environment to test the vehicle lane-changing strategy using highway scenarios, and use two statistics, the mean absolute error and the mean absolute relative error, to perform error statistics on the model.

Citation Information

Patent Citations

  • Large commercial vehicle lane change decision-making method based on deep learning

    CN113954837A