Automatic lane changing safety reinforcement learning control method giving consideration to expected safety of rear vehicle
By introducing a rear-vehicle collision time evaluation mechanism and reinforced learning architecture in autonomous driving technology, the problem of lane change safety and traffic flow stability in dynamic traffic environments is solved, and safer and more efficient automatic lane change control is achieved.
Patent Information
- Application Number
- CN202510500340.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-06-27
AI Technical Summary
In a dynamic traffic environment, it is difficult for existing autonomous driving technology to take into account lane change safety and traffic flow stability, especially to ignore the safety impact of the rear vehicle, which may lead to the rear vehicle being decelerated rapidly, affecting traffic flow and increasing the risk of accidents.
A reinforcement learning control method for automatic lane change safety taking into account the expected safety of the rear vehicle is proposed. By generating the motion state information of the bicycle and the rear vehicle based on the input information of the vehicle sensor, combining Lagrangian constraint optimization and a double-layer deep Q network, a reinforcement learning architecture is built, the collision time of the rear vehicle is evaluated and the safety cost weight is dynamically adjusted, and the lane change decision strategy is optimized.
This method uses real-time perception of vehicle environment information, evaluates the emergency deceleration risk of the vehicle afterwards, avoids safety hazards during lane change, and ensures that the bicycle can apply better lane change control strategies in a dynamic traffic environment, taking into account lane change safety and traffic flow efficiency.
Smart Images

Figure CN120207333A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving of intelligent connected vehicles, and in particular to an automatic lane change safety reinforcement learning control method that takes into account the expected safety of the following vehicle, which is applicable to the autonomous lane change decision-making and control of vehicles in a multi-lane highway scenario, aiming to solve the problem that it is difficult to balance lane change safety and traffic flow stability in a dynamic traffic environment. Background Technique
[0002] With the development of autonomous driving technology, the application of reinforcement learning in vehicle lane change decision-making control has gradually become a research hotspot. Traditional lane change decision-making methods cannot adapt to complex and changeable traffic scenarios, or have insufficient real-time performance due to high computational complexity. Existing reinforcement learning methods often focus on the assessment of the time to collision with the leading vehicle in lane change decisions, ignoring the safety impact of the following vehicle, which may cause the following vehicle to be forced to decelerate sharply during the lane change process, thus affecting traffic flow and increasing the potential risk of traffic accidents.
[0003] Therefore, how to construct an automatic lane change control method that takes into account both safety and efficiency by combining the expected safety of the following vehicle and driving efficiency under the reinforcement learning framework has become a key issue in the development of autonomous driving technology. For this reason, the present invention proposes an automatic lane change safety reinforcement learning control method that takes into account the expected safety of the following vehicle, and solves the problem that it is difficult to balance lane change safety and traffic flow stability in a dynamic traffic environment by taking into account the time to collision with the following vehicle. Summary of the Invention
[0004] The objective of the present invention is to provide an automatic lane change safety reinforcement learning control method that takes into account the expected safety of the following vehicle, which is mainly applied to the autonomous lane change decision-making and control of vehicles in a multi-lane highway environment.
[0005] An automatic lane change safety reinforcement learning control method that takes into account the expected safety of the following vehicle includes the following steps:
[0006] S1. Based on the input information of vehicle sensors, generate the motion state information of the host vehicle and the following vehicle as the reinforcement learning state space;
[0007] The reinforcement learning state space is represented by a high-dimensional state vector, including the presence state, position information, and speed information of surrounding target vehicles, so as to ensure that the reinforcement learning agent can comprehensively perceive the dynamic information of the host vehicle and the surrounding environment and provide a basis for optimizing the lane change decision;
[0008] S2. Based on Lagrangian constraint optimization, fuse a double-layer deep Q network to establish a Lagrangian safety reinforcement learning architecture;
[0009] A two-layer deep Q network structure is used, in which the online network is used for real-time strategy evaluation and update, and the target network is used to smooth the Q value estimation to reduce the estimation deviation caused by state transfer;
[0010] The Lagrangian constraint method is introduced, and the adaptive multiplier mechanism is used to dynamically adjust the safety cost weight according to the real-time traffic environment and vehicle status information to achieve adaptive optimization of safety constraints.
[0011] Combining the above strategy update with constraint optimization to form a joint optimization model enables the agent to dynamically balance speed and safety while maintaining efficient cruising, avoiding the reduction of driving efficiency due to over-conservatism;
[0012] S3. Comprehensively consider speed rewards, overtaking rewards, collision penalties, and the spatiotemporal safety of the front and rear vehicles, construct a reinforcement learning reward function and a safety cost function, and obtain an automatic cruise reinforcement learning control agent after training;
[0013] By evaluating the difference between the current speed of the vehicle and the target speed, a speed reward mechanism is set; positive rewards are triggered when overtaking is successfully completed; negative rewards are triggered when a collision occurs; based on the collision time and space margin, time and space safety constraints are set to ensure safety during lane changes. By comprehensively considering the above factors, a reinforcement learning reward function and a safety cost function are constructed to train the automatic cruise reinforcement learning control agent.
[0014] Preferably, in step S1, the state space is represented by a high-dimensional state vector, including the existence status, position information and speed information of surrounding target vehicles, so as to ensure that the reinforcement learning agent can effectively perceive environmental information and optimize decisions under different traffic conditions.
[0015] Preferably, in step S2, the Lagrangian constrained optimization method is introduced to construct a joint optimization framework, soft constraint control of the safety cost function is achieved by setting multipliers, and strategy learning is performed with the following optimization objectives:
[0016] Subject to
[0017] Among them, π represents the reinforcement learning strategy, that is, the behavior rules adopted by the agent in each state, represents the expectation operator, which is used to perform weighted summation of all state transition paths, r t (s t ,a t ) is the state s of the agent at time step t t And perform action a t The immediate reward obtained, γ is the discount factor used to control the influence weight of future rewards in the current decision, ct (s t , a t ) is the safety cost function for the corresponding state-action pair, which is used to measure the impact of this behavior on the overall traffic safety. δ is the maximum tolerable safety cost threshold set by the system. While ensuring that the agent maximizes the cumulative reward, this optimization objective restricts the actions taken by the agent so that the safety cost of the overall system does not exceed the specified upper limit, thereby achieving a dynamic balance between efficiency and safety.
[0018] Preferably, in step S2, this constrained problem is transformed into a Lagrangian form:
[0019]
[0020] Preferably, in step S2, a double-deep Q-network structure is used to train the reinforcement learning agent. The online network is used for real-time policy evaluation and action selection, while the target network performs soft parameter updates at fixed intervals to reduce the Q-value estimation bias and enhance the stability of policy training. Its Q-value update function is:
[0021] Q(s t , a t ) ← Q(s t , a t ) + α[r t + γ max a' Q'(s t+1 , a') - Q(s t , a t )]
[0022] where Q(s t , a t ) is the estimated value of the current state-action pair, α is the learning rate, which controls the amplitude of Q-value update, and Q'(s t+1 , a') is the estimated value of each action a' in the next state s t+1 . max a 'Q'(·) represents selecting the future action with the largest estimated value. Through this double-network structure, this function effectively alleviates the problem of overestimation of Q-values in traditional Q-learning and improves the stability and reliability of the training process.
[0023] Preferably, in step S2, the Lagrange multiplier λ is updated iteratively in the following form to achieve adaptive optimization of the safety cost weight:
[0024]
[0025] where β is the learning rate for updating the Lagrange multiplier, which controls its adjustment speed, Let \(C_{avg}\) be the average safety cost of the state-action pair under the current policy, and \(\delta\) represents the upper threshold of the safety cost. If the safety cost caused by the current policy is higher than the preset threshold, then \(\lambda\) increases, thereby enhancing the constraint on safety. Conversely, if the current policy has good safety, then \(\lambda\) will gradually decrease, thereby increasing the efficiency-oriented weight of the policy.
[0026] Preferably, in step S3, the rear vehicle spatio-temporal safety reward term in the reward function is calculated based on the time to collision, where the time to collision is obtained from the ratio of the longitudinal distance between the host vehicle and the rear vehicle to the speed of the rear vehicle. When the lane change of the host vehicle causes the time to collision to be lower than the preset safety threshold, the reinforcement learning agent will receive a negative reward and adjust the decision-making strategy to maintain a reasonable safety margin and reduce the impact on the safety of the rear vehicle.
[0027] Preferably, a safety cost function is introduced in step S3 to construct a multi-dimensional safety index for restricting unsafe lane change behaviors; and through the joint optimization objective, the weight ratio of this cost function in the overall return is dynamically adjusted during the reinforcement learning process to improve the adaptability of the model to complex traffic environments.
[0028] Advantages of the present invention: This method constructs a lane change decision-making model by real-time sensing vehicle environment information, introduces a rear vehicle collision time evaluation mechanism to evaluate the sudden deceleration risk of the rear vehicle, thereby avoiding potential safety hazards during the lane change process. At the same time, a safety reinforcement learning algorithm is used to optimize the lane change decision-making strategy to ensure that the host vehicle can apply a better lane change control strategy in a dynamic traffic environment. This method can balance lane change safety and traffic flow efficiency and has significant engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic diagram of the principle of an automatic lane change safety reinforcement learning control method that takes into account the expected safety of the rear vehicle according to the present invention;
[0030] Figure 2 is a convergence curve graph of the double-layer deep Q-network policy in an automatic lane change safety reinforcement learning control method that takes into account the expected safety of the rear vehicle according to the present invention;
[0031] Figure 3 is a comparison curve graph of vehicle acceleration and time to collision of an automatic lane change safety reinforcement learning control method that takes into account the expected safety of the rear vehicle. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.
[0033] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.
[0034] In the present invention, words such as "first", "second" and the like do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0035] Embodiment 1
[0036] As Figure 1 shown, an automatic lane-changing safety reinforcement learning control method that takes into account the expected safety of the following vehicle includes the following steps:
[0037] S1. Based on the input information of vehicle sensors, generate the motion state information of the host vehicle and the following vehicle as the reinforcement learning state space.
[0038] Preprocess the input information collected by the vehicle sensors, fuse the position, speed, and acceleration information of the surrounding vehicles with the state information of the host vehicle to form a state space for reinforcement learning, which is used as the input for subsequent decision optimization.
[0039] The reinforcement learning state space is represented by a high-dimensional state vector to comprehensively describe the dynamic information of the host vehicle and the surrounding environment. The state vector is defined as:
[0040] S = {p res , x, y, v x , v y}
[0041] where p res represents the presence state of the surrounding target vehicle, x and y are the position information of the surrounding vehicles, and v x , v y are the speed information of the surrounding vehicles. Through this state representation, the reinforcement learning agent can effectively perceive the environmental information and optimize decisions under different traffic conditions.
[0042] S2. Based on Lagrangian constraint optimization, fuse the double-layer deep Q network to establish a Lagrangian safety reinforcement learning architecture.
[0043] Express the constraint problem in reinforcement learning as a Lagrangian optimization form, the goal of which is to obtain the maximum cumulative reward while constraining the behavior of the agent to meet the preset safety standards. The Lagrangian form of this optimization problem is defined as follows:
[0044]
[0045] where: π represents the policy of the agent; λ is the Lagrange multiplier used to control the weight of the safety cost; γ ∈ (0, 1) is the discount factor, indicating the attenuation degree of future rewards; r t (s t , a t ) is the immediate reward obtained at time step t for state s t and action a t ; λc t (s t , a t ) is the safety cost function at time step t, measuring the impact of this action on system safety; represents the expectation of the entire policy execution process.
[0046] The reinforcement learning agent is trained using a double - layer deep Q - network architecture, which includes an online network and a target network. The online network is used to update the Q - value in real - time and optimize the lane - changing policy. The target network is used to calculate the target Q - value and reduce the estimation bias during the training process through parameter synchronization at fixed intervals. The update rule of the Q - value function is as follows:
[0047] Q(s t , a t ) ← Q(s t , a t ) + α[r t + γmax a' Q'(s t+1 , a') - Q(s t , a t )]
[0048] where: Q(s t , a t ) is the Q - value estimate of the current state - action pair; α ∈ (0, 1) is the learning rate; r t is the current reward; γ is the discount factor; Q'(s t+1 , a') is the Q - value estimate of the target network for each action a' at the next state s t+1 ; max a' Q'(s t+1 , a') represents selecting the action with the maximum Q - value as the expected return.
[0049] Compared with the traditional deep Q-network method, the double-layer deep Q-network can effectively reduce the overestimation problem of Q-values, improve the training stability, and reduce the unreasonable lane-changing decisions caused by the Q-value estimation deviation. During the reinforcement learning training process, an experience replay mechanism is adopted to store historical decision-making data to improve the sample utilization rate. Set the experience replay pool D, and the agent randomly extracts a small batch of samples from the replay pool for training during training, and optimizes the loss function:
[0050]
[0051] Among them, y is the target Q-value, which is calculated by the target network:
[0052] y = r + γQ(s', argmax a' Q(s', a'; θ); θ - )
[0053] Among them, r is the immediate reward, γ is the discount factor, s' is the next state, a' is the next action, θ is the online network parameter, and θ - is the target network parameter.
[0054] The target network parameters are synchronized at fixed intervals:
[0055] θ - ← θ
[0056] Thereby reducing the oscillation during training, improving the convergence performance of reinforcement learning, and enabling the agent to achieve stable and optimized lane-changing decisions in a complex traffic environment.
[0057] To achieve dynamic adjustment of safety, an adaptive update mechanism of Lagrange multipliers is introduced. Its update formula is as follows:
[0058]
[0059] Among them: λ is the current Lagrange multiplier; β is the learning rate; is the expected safety cost under the current policy; δ is the preset maximum acceptable safety cost threshold.
[0060] This mechanism dynamically adjusts the safety cost weight in each round of training, so that when the safety cost exceeds the threshold δ, the constraint intensity is increased; while when the safety is sufficient, the constraint is weakened and the behavior freedom of the agent is enhanced.
[0061] Through the above structure, the constructed joint optimization model can effectively improve the flexibility and efficiency of lane-changing decisions while ensuring the safe behavior of the agent, and avoid the problem of reduced traffic efficiency caused by excessive conservatism in traditional methods.
[0062] S3. Considering speed rewards, overtaking rewards, collision penalties, and the spatio-temporal safety of the leading vehicle and the following vehicle comprehensively, construct a reinforcement learning reward function and a safety cost function, and obtain an adaptive cruise reinforcement learning control agent after training.
[0063] When constructing the reward function, comprehensively consider the driving safety and traffic fluency of the vehicle to achieve safe and efficient lane-changing control. The reward function R consists of speed rewards, overtaking rewards, collision penalties, and spatio-temporal safety constraints, and is defined as follows:
[0064] R = ω1R speed + ω2R over - ω3R coll + ω4R safety
[0065] Among them, ω1, ω2, ω3, and ω4 are the weight parameters of each reward item, which are optimized and adjusted through experiments to ensure that the reward function has good numerical stability and policy optimization ability during the training process.
[0066] Speed reward R speed It is used to encourage the agent to approach and maintain the target cruise speed, and the calculation method is as follows:
[0067]
[0068] Among them, v is the current vehicle speed. When the vehicle speed reaches the preset value, the reward value reaches the maximum; when it is less than a certain range, no reward is given to ensure that the vehicle stably maintains within a reasonable cruise speed range.
[0069] Overtaking reward R over A positive reward is given when the agent successfully completes a safe overtaking to optimize the driving path and improve traffic efficiency.
[0070] Collision penalty R coll A high-weight negative reward is imposed when a collision event is detected to ensure that the agent always prioritizes avoiding dangerous behaviors during training.
[0071] Spatio-temporal safety reward R safety The collision time index is used to evaluate the safety distance between the vehicle and surrounding targets, and a negative reward is imposed when the safety margin is insufficient to prompt the agent to actively adjust the strategy to maintain a reasonable safety distance. The calculation method is as follows:
[0072]
[0073] Among them, Δd is the distance from the nearest vehicle behind the host vehicle to the host vehicle in the target lane where the host vehicle changes lanes, v *It refers to the speed of the vehicle closest to the ego vehicle in the target lane for the ego vehicle to change lanes and behind the ego vehicle. When the TTC is less than the preset value, the negative reward increases to prompt the agent to take collision avoidance measures and ensure driving safety.
[0074] When constructing the safety cost function, comprehensively consider the potential traffic safety risks that may occur during the automatic lane change process. The safety cost function consists of the following three parts: time-to-collision cost, vehicle distance cost, and collision event cost.
[0075] Among them, the time-to-collision cost is dynamically calculated based on the time-to-collision between the vehicle and the target in front. When the time-to-collision value is lower than the preset safety threshold, the cost decreases as the time-to-collision decreases;
[0076] The vehicle distance cost is calculated based on the difference between the actual following distance and the safety threshold. When the distance is lower than the threshold, the cost increases.
[0077] The collision event cost is a fixed high-weight cost, which is only triggered when a vehicle collision occurs.
[0078] In the traditional reinforcement learning lane change strategy, usually only the time-to-collision of the vehicle in front is considered, that is, the safety of lane change is judged by calculating the TTC between the ego vehicle and the target vehicle in front. The present invention further introduces the TTC of the vehicle behind for evaluation to ensure that the ego vehicle's lane change will not pose too much safety risk to the vehicle behind. When the ego vehicle's lane change causes the TTC of the vehicle behind to be lower than the safety threshold, the agent will also receive a negative reward to prompt it to adjust the lane change strategy and avoid causing the vehicle behind to brake suddenly due to lane change, thereby improving the safety of lane change.
[0079] In the method of the present invention, the calculation of the TTC of the vehicle behind is based on the distance from the closest vehicle behind the ego vehicle to the ego vehicle in the target lane for the ego vehicle to change lanes and the driving speed of this vehicle. When the TTC of the vehicle behind is lower than the set threshold, the agent will adjust the lane change strategy to maintain a reasonable safety margin and reduce the interference to the safety of the vehicle behind. In addition, the TTC of the vehicle in front is still used as part of the safety constraint to ensure that there is no forward collision risk during the lane change process.
[0080] As Figure 2 shown, during the reinforcement learning training process, the reward value based on the double-layer deep Q-network method gradually increases with the number of training rounds and finally stabilizes, indicating that the agent can effectively learn a reasonable lane change strategy. By optimizing the reward function, this method fully considers the time-to-collision of the vehicle behind, enabling the agent to actively adjust its decision during the training process to reduce the impact on the driving state of the vehicle behind and ensure the safety and smoothness of the lane change process. The training results show that under the constraint of the TTC of the vehicle behind, the agent can learn to perform lane change operations at the appropriate time, thereby improving the overall driving safety and traffic flow stability.
[0081] To evaluate the trained reinforcement learning agent and analyze its lane-changing effect, a fixed test scenario was constructed for repeated testing. In this environment, a comparative analysis was conducted between the double-layer deep Q-network method without considering the following vehicle collision time constraint and the double-layer deep Q-network method considering the following vehicle collision time. The experimental results show that the double-layer deep Q-network method without considering the following vehicle collision time may cause the following vehicle to decelerate sharply in some cases, affecting driving safety, while the double-layer deep Q-network method considering the following vehicle collision time can effectively reduce the interference to the driving state of the following vehicle, make the lane-changing process smoother, and improve the stability of the overall traffic flow. The specific results are as follows.
[0082] As Figure 3 shown, in the fixed test scenario, the lane-changing control strategy based on the double-layer deep Q-network method has different effects on the driving state of the following vehicle. The double-layer deep Q-network method without considering the following vehicle collision time may cause the following vehicle to decelerate sharply during the lane-changing process, resulting in a sharp fluctuation in the acceleration of the following vehicle and a decrease in the collision time within a short period, thus increasing the potential collision risk. In contrast, the double-layer deep Q-network method considering the following vehicle collision time can adjust the control strategy more smoothly during the lane-changing process, avoid having too much impact on the driving state of the following vehicle, ensure that the collision time remains within a reasonable range, and thus improve the overall driving safety and traffic flow stability.
[0083] To verify the effectiveness of the method of the present invention, a comparative analysis was conducted on the speed performance in the automatic cruise task with and without considering the following vehicle collision time. The experiments were evaluated from three key indicators: traffic flow speed, the average speed of the host vehicle, and the average deceleration of the following vehicle. The specific performance comparisons are shown in Tables 1 and 2
[0084] Table 1 Comparison of speed performance considering the following vehicle collision time under different densities
[0085]
[0086] Table 2 Comparison of speed performance without considering the following vehicle collision time under different densities
[0087]
[0088]
[0089] In the method of using the time-to-collision evaluation mechanism in low-density, medium-density, and high-density traffic scenarios, the average speed of the ego vehicle is 23.83 m / s, 23.61 m / s, and 23.27 m / s respectively. Compared with the strategy that does not consider the time-to-collision of the following vehicle, the speed reduction of the ego vehicle is controlled within the range of 0.08% to 0.17%, indicating that the traffic efficiency is not significantly affected. At the same time, the average speed of the following vehicle is optimized from 1.65 m / s to 1.62 m / s in the high-density scenario, and the speed fluctuation range of the following vehicle in the low-density scenario is reduced to 1.51 m / s to 1.62 m / s, proving that the interference of the ego vehicle's lane-changing behavior to the following vehicle is effectively suppressed by the time-to-collision constraint mechanism. Especially when the traffic density increases, in the strategy that does not use the time-to-collision evaluation of the following vehicle, the speed of the following vehicle shows an abnormal upward trend with the increase of density, presumably caused by the frequent and aggressive lane-changing of the ego vehicle forcing the following vehicle to accelerate and avoid; while the present invention makes the speed fluctuation of the following vehicle tend to be stable through the dynamic penalty mechanism, and the speed difference between the ego vehicle and the traffic flow is stably maintained in the range of 1.05 m / s to 1.89 m / s, verifying the robustness of this method in complex scenarios.
[0090] The present invention comprehensively verifies the adaptability of the method in sparse traffic flow, normal traffic flow, and high-density congestion scenarios by introducing different density scenario tests. The low-density scenario focuses on efficiency optimization, the medium-density scenario balances safety and efficiency, and the high-density scenario needs to cope with the challenges of frequent interactions and narrow vehicle distances. The experimental results show that in the high-density scenario, the average speed fluctuation of the following vehicle is reduced by 4.2%, and the speed of the ego vehicle is only reduced by 0.17%, proving that the present invention can adaptively adjust the strategy through the dual value estimation mechanism of the double deep Q network, avoiding decision failure caused by traffic state changes.
[0091] The core of the present invention is to incorporate the quantitative evaluation of the time-to-collision of the following vehicle into the reinforcement learning framework. When the time-to-collision of the following vehicle is lower than the 1.2-second safety threshold due to the ego vehicle's lane-changing, the reinforcement learning agent will be driven by negative rewards and autonomously adjust the lane-changing timing and acceleration to avoid causing emergency braking of the following vehicle. Compared with the traditional reinforcement learning method that only focuses on the safety of the leading vehicle, the present invention constructs a decision-making model with social compatibility through a multi-dimensional time-to-collision penalty mechanism, providing an innovative solution for the lane-changing control of autonomous vehicles in a dynamic traffic environment.
[0092] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A safety reinforcement learning control method for automatic lane change taking into account the expected safety of the following vehicle, characterized in that: The following steps are involved: S1, based on the vehicle sensor input information, generate the motion state information of the vehicle and the following vehicle as the reinforcement learning state space; The reinforcement learning state space is represented by a high-dimensional state vector, including the existence status, position information and speed information of surrounding target vehicles, so as to ensure that the reinforcement learning agent can fully perceive the dynamic information of the vehicle and the surrounding environment, providing a basis for optimizing lane change decisions; S2. Based on Lagrangian constraint optimization, a two-layer deep Q network is integrated to establish a Lagrangian secure reinforcement learning architecture; A two-layer deep Q network structure is used, in which the online network is used for real-time strategy evaluation and update, and the target network is used to smooth the Q value estimation to reduce the estimation deviation caused by state transfer; The Lagrangian constraint method is introduced, and the adaptive multiplier mechanism is used to dynamically adjust the safety cost weight according to the real-time traffic environment and vehicle status information to achieve adaptive optimization of safety constraints. Combining the above strategy update with constraint optimization to form a joint optimization model enables the agent to dynamically balance speed and safety while maintaining efficient cruising, avoiding the reduction of driving efficiency due to over-conservatism; S3. Comprehensively consider speed rewards, overtaking rewards, collision penalties, and the spatiotemporal safety of the front and rear vehicles, construct a reinforcement learning reward function and a safety cost function, and obtain an automatic cruise reinforcement learning control agent after training; By evaluating the difference between the current speed of the vehicle and the target speed, a speed reward mechanism is set; positive rewards are triggered when overtaking is successfully completed; negative rewards are triggered when a collision occurs; based on the collision time and space margin, time and space safety constraints are set to ensure safety during lane changes. By comprehensively considering the above factors, a reinforcement learning reward function and a safety cost function are constructed to train the automatic cruise reinforcement learning control agent.
2. The automatic lane change safety reinforcement learning control method taking into account the expected safety of the following vehicle according to claim 1, characterized in that: In step S1, the state space is represented by a high-dimensional state vector, including the existence status, position information and speed information of surrounding target vehicles, so as to ensure that the reinforcement learning agent can effectively perceive environmental information and optimize decisions under different traffic conditions.
3. The method according to claim 1, characterized in that In step S2, the Lagrangian constrained optimization method is introduced to construct a joint optimization framework, and the soft constraint control of the safety cost function is realized by setting the multiplier, and the strategy learning is performed with the following optimization objectives: Among them, π represents the reinforcement learning strategy, that is, the behavior rules adopted by the agent in each state, represents the expectation operator, which is used to perform weighted summation of all state transition paths, r t (s t ,a t ) is the state s of the agent at time step t t And perform action a t The immediate reward obtained, γ is the discount factor used to control the influence weight of future rewards in the current decision, c t (s t ,a t ) is the safety cost function of the corresponding state-action pair, which is used to measure the impact of the behavior on the overall traffic safety, and δ is the maximum tolerable safety cost threshold set by the system. This optimization goal ensures that the agent maximizes the cumulative reward while constraining the behavior it takes so that the safety cost of the overall system cannot exceed the specified upper limit, thereby achieving a dynamic balance between efficiency and safety.
4. The method according to claim 1, characterized in that In step S2, the constraint problem is transformed into Lagrangian form:
5. The method according to claim 1, characterized in that In step S2, a two-layer deep Q network structure is used to train the reinforcement learning agent, where the online network is used for real-time strategy evaluation and behavior selection, and the target network performs soft parameter updates at fixed intervals to reduce the Q value estimation bias and enhance the stability of strategy training. The Q value update function is: Q(s t ,a t )←Q(s t ,a t )+α[r t +γmax a' Q'(s t+1 ,a')-Q(s t ,a t )] Among them, Q(s t ,a t ) is the estimated value of the current state action pair, α is the learning rate, which controls the Q value update amplitude, Q'(s t+1 ,a') is the target network in the next state s t+1 The valuation of each action a', max a' Q'(·) represents the selection of the future action with the largest valuation. This function effectively alleviates the problem of overestimation of Q value in traditional Q learning through the dual network structure, and improves the stability and reliability of the training process.
6. The method according to claim 1, characterized in that In step S2, the update of the Lagrange multiplier λ is iteratively performed in the following form to achieve adaptive optimization of the safety cost weight: Among them, β is the learning rate of the Lagrange multiplier update, which controls its adjustment speed. is the average security cost of the state-action pair under the current strategy, and δ represents the upper threshold of the security cost. If the security cost caused by the current strategy is higher than the preset threshold, λ will increase, thereby improving the constraint on security. On the contrary, if the current strategy has good security, λ will gradually decrease, thereby improving the efficiency-oriented weight of the strategy.
7. The method according to claim 1, characterized in that In step S3, the spatiotemporal safety reward term of the following vehicle in the reward function is calculated based on the collision time, where the collision time is derived from the ratio of the longitudinal distance between the ego vehicle and the following vehicle and the speed of the following vehicle. When the collision time caused by the lane change of the ego vehicle is lower than the preset safety threshold, the reinforcement learning agent will receive a negative reward and adjust the decision-making strategy to maintain a reasonable safety margin and reduce the impact on the safety of the following vehicle.
8. The method according to claim 1, characterized in that: In step S3, a safety cost function is introduced to construct a multi-dimensional safety index to limit unsafe lane-changing behaviors. And through the joint optimization goal, the weight ratio of the cost function in the overall reward is dynamically adjusted during the reinforcement learning process to improve the adaptability of the model to complex traffic environments.
Citation Information
Cited By
Automatic driving lane changing decision-making method fusing safety integrity framework and reinforcement learning
CN121291489A
An Autonomous Driving Lane Changing Decision-Making Method Integrating Safety Integrity Framework and Reinforcement Learning
CN121291489B