Autonomous Driving Decision-making Method Based on Security-Enhanced Multi-Agent Reinforcement Learning
By introducing autonomous driving decision-making methods that enhance the strengthening learning of multiple agiles, using the safety awareness module and multi-agile decision-making module, the safety and efficiency problems of autonomous vehicles in the highway convergence area are solved, and the coordinated decision-making of multiple vehicles and safe and efficient inclusion are achieved.
Patent Information
- Application Number
- CN202510712141.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing autonomous driving methods are difficult to achieve multi-vehicle collaborative decision-making in highway confluence areas, which pose safety risks and are inefficient, especially in hybrid traffic environments, and are difficult to cope with the unpredictable behavior of artificially driven vehicles.
Adopt autonomous driving decision-making method based on safety-enhanced multi-agent reinforcement learning, and by introducing safety awareness modules and multi-agent decision-making modules, combining Bayesian reasoning and action shielding technology, a probability model is constructed to evaluate collision risks, and multi-vehicle collaboration is realized through shared network structure.
It improves the safety and efficiency of autonomous vehicles in high-speed inlet zones, and can make accurate safety decisions in complex traffic environments, reduce collision rates and improve the organizational efficiency of traffic flow.
Smart Images

Figure CN120235061B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to an autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning, which is applicable to multi-vehicle cooperative control and safe obstacle avoidance in the merging area of an expressway under a mixed traffic environment. Background Art
[0002] With the continuous progress of autonomous driving technology and its widespread integration into actual traffic scenarios, the merging area of an expressway remains one of the major challenges for autonomous driving, especially in a mixed traffic environment, where its challenges become even more prominent. In a mixed traffic environment, autonomous vehicles (AVs) and human-driven vehicles (HDVs) travel together, and the interaction between vehicles becomes extremely complex, and the safety risks are also significantly increased. In the merging area of an expressway, it is necessary to quickly and accurately adjust the vehicle speed and lane position to fit the rapidly changing traffic conditions, which undoubtedly greatly increases the complexity and operation difficulty of the autonomous driving system.
[0003] In terms of safety issues, existing autonomous driving methods have many shortcomings. Facing the dynamic interaction and sudden changes in the high-speed merging area, these methods are difficult to accurately predict the behavior of human-driven vehicles. Traditional rule-based control methods usually rely on pre-set static thresholds to guide vehicle operations. However, in the merging area, the vehicle speed, position, and relative distance are all changing rapidly, and static thresholds cannot adapt to these dynamic changes in a timely manner, resulting in possible deviations in vehicle decision-making and increasing the collision risk. Currently, many related studies overly focus on the optimization of traffic efficiency and to a certain extent ignore the high collision risk existing in the merging area itself, so that in practical applications, potential safety hazards have not been effectively resolved.
[0004] Each AVs reinforcement learning method regards other vehicles only as part of the environment and lacks the ability to make collaborative decisions with other vehicles, which makes the efficiency extremely low in scenarios of multi-vehicle interaction. Even if some methods attempt to combine model predictive control with reinforcement learning, due to the limitations of its per-AVs architecture, it is still difficult to meet the actual needs of multi-vehicle cooperation. When facing the unpredictable behaviors of human-driven vehicles, existing algorithms perform unsatisfactorily and it is difficult to balance traffic efficiency while ensuring safety. In addition, existing perception systems and behavior prediction algorithms have great deficiencies in capturing the complex behavior patterns of human-driven vehicles, resulting in a significant reduction in the accuracy of predicting the intentions and behaviors of surrounding vehicles.
[0005] Traditional decision-making and control methods are mainly based on fixed rules and models and lack the flexibility and adaptive ability to cope with changing traffic environments. Most existing autonomous driving systems operate independently, and there is a lack of effective collaborative decision-making and information sharing mechanisms between vehicles, making it difficult to achieve efficient organization and optimization of the entire traffic flow and unable to give full play to the advantages of autonomous driving technology. Summary of the Invention
[0006] The object of the present invention is to propose an autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning. By introducing a safety awareness and multi-agent decision-making module, it is possible to achieve safe and efficient merging of AVs in the highway merging area.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] An autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning includes the following steps:
[0009] Step 1. Obtain environmental observation information including the current vehicle state information, observable surrounding vehicle information, and environmental information. Perform data preprocessing and integration on the obtained environmental observation information to construct a state space and an action space.
[0010] Step 2. Design a safety awareness module based on security-enhanced multi-agent reinforcement learning.
[0011] First, construct a probability model to model the time to collision (TTC) and the blind spot threat (HZT) respectively. Adopt a piecewise likelihood function combined with Gaussian and uniform distributions to flexibly capture environmental uncertainties.
[0012] For different safety levels, construct different forms of likelihood functions respectively. Update the probability distribution of the vehicle safety level through Bayes' formula. Finally, determine the safety level of AVs by maximizing the posterior probability.
[0013] Step 3. Design a multi-agent decision-making module to achieve multi-vehicle cooperation.
[0014] Define a state space, an action space, and a reward function.
[0015] The state vector of each AV is used to comprehensively describe the surrounding environment. The action space includes 6 discrete actions. The reward function includes a safety reward, a collision reward, a speed reward, a time headway reward, and a lane change reward.
[0016] Construct a shared policy network structure, including an actor network and a critic network. The two share the same low-level representation, including position state and speed state, to optimize the policy gradient update and value function learning process.
[0017] Step 4. Introduce an action masking module. By predefined rules, mask invalid or unsafe actions and restrict the output of the multi-agent decision-making module to ensure that the system performs safe operations.
[0018] The present invention has the following advantages:
[0019] As described above, the present invention relates to an autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning. Among them, the security awareness module introduced in the present invention evaluates the vehicle collision risk through Bayesian inference, constructs a probability model to model key indicators, and adopts a piecewise likelihood function combined with Gaussian and uniform distributions to flexibly capture environmental uncertainties. This risk assessment method improves the understanding ability of AVs for complex traffic environments and helps the vehicle make more accurate safety decisions in the high-speed merging area. Secondly, the present invention adopts a multi-agent decision-making module, which is based on the multi-agent reinforcement learning algorithm of A2C. By introducing the advantage function and combining the policy gradient and value function methods, the learning efficiency and stability are improved. The multi-agent decision-making module realizes the efficient cooperation of multiple AVs in complex traffic environments, ensures that the vehicle can successfully complete the merging task, and improves traffic efficiency. In addition, the multi-agent decision-making module defines the state space, action space, and reward function to ensure that the agent executes tasks safely and efficiently in complex traffic environments. In addition, the present invention also introduces an action masking technology. According to the security assessment results, unsafe actions are masked, unreasonable or high-risk driving behaviors are restricted, and vehicle safety is ensured. Through predefined action masking rules, the logits of unsafe actions are set to very small negative values, making the probabilities of these actions after the Softmax function close to 0, thereby ensuring that the vehicle only performs safe operations. This action masking mechanism effectively avoids the vehicle from performing dangerous actions in the high-speed merging area and further improves driving safety. In terms of the network structure, the actor network and critic network proposed in the present invention share the same low-level representation, including position state and speed state. These states are input into a separate fully connected layer for processing, and then combined and passed to the final fully connected layer. The output results are used for the actor network and critic network. By sharing the network structure, the parameter utilization efficiency and learning speed of the model are improved, enabling the model to adapt to complex traffic environments faster. In the training and evaluation phase, the present invention implements the high-speed merging area scenario in the Gym-based highway-env simulator. By setting different traffic densities, the curriculum learning method is used for training. After training is completed, the performance of the vehicle in various scenarios is evaluated, including safety, efficiency, and stability. Specifically, training starts from a simple scenario with low traffic density and gradually transitions to a complex scenario with high traffic density to improve the adaptability and generalization ability of the model. Through comparative experiments with existing methods, the performance of the method proposed in the present invention in different traffic modes is verified. The results show that the method of the present invention has significant advantages in terms of safety, efficiency, and stability, can maintain a low collision rate under different traffic densities, and at the same time maintain a high average speed and small speed fluctuations. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1Flowchart of the autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning in the embodiments of the present invention;
[0021] Figure 2 Network structure diagram of the security awareness module in the embodiments of the present invention;
[0022] Figure 3 Block diagram of the autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning in the embodiments of the present invention;
[0023] Figure 4 Architecture diagram of the action masking module in the embodiments of the present invention. Detailed implementation manners
[0024] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners:
[0025] Embodiment
[0026] In this embodiment, an autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning is described. In this method, a security awareness module and a multi-agent decision-making module are designed. By introducing Bayesian inference and action masking technologies, the safety of vehicles in the high-speed merging area is improved. At the same time, a deep reinforcement learning algorithm is used to achieve multi-vehicle cooperation, improving the merging efficiency. Specifically, the present invention designs a security awareness module and a multi-agent decision-making module. By introducing Bayesian inference and action masking technologies, the safety of vehicles in the high-speed merging area is improved while a deep reinforcement learning algorithm is used to achieve multi-vehicle cooperation, improving the merging efficiency. The security awareness module models the time to collision (TTC) and the hazard in the blind zone (HZT) by constructing a probability model, and uses a piecewise likelihood function combined with Gaussian and uniform distributions to flexibly capture environmental uncertainties, improving the comprehensiveness and reliability of safety analysis. The multi-agent decision-making module defines the state space, action space, and reward function, and dynamically adjusts the reward value to solve the dynamic interaction problem in the highway merging area. In addition, the present invention also constructs a shared network structure, including an actor network and a critic network, optimizing the policy gradient update and value function learning process. In addition, by introducing an action masking module, unsafe actions are masked to ensure that the system selects safe operations. The experimental environment scenario is set in the high-speed merging area and the highway-env simulator to evaluate the performance of the vehicle under different traffic densities, ensuring that the vehicle can successfully complete the merging task and improving the obstacle avoidance ability and driving efficiency. The present invention significantly improves the safety, efficiency, and adaptability of AVs in the highway merging area by combining the security awareness module, multi-agent decision-making module, shared network structure, and action masking module.
[0027] As Figure 1 shown, the autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning includes the following steps:
[0028] Step 1. Obtain environmental observation information including the current vehicle state information, observable surrounding vehicle information, and environmental information. Perform data preprocessing and integration on the obtained environmental observation information to construct a state space and an action space.
[0029] In this embodiment, the AVs obtain environmental observation information in the highway-env simulation environment. The highway-env simulation environment is a simulation environment based on the Gym library, specifically designed for decision-making of AVs in highway scenarios.
[0030] The process of obtaining environmental observation information is as follows:
[0031] I. Acquisition of vehicle state information: For each autonomous vehicle (AVs), obtain its state information, including position, speed, and acceleration.
[0032] II. Acquisition of observable surrounding vehicle information: Obtain the position and speed information of human-driven vehicles (HDVs) and AVs within the observable range around each AVs to construct the observation state of the vehicle.
[0033] III. Acquisition of environmental information: Obtain environmental information of the road, including topological structure, traffic signs, traffic lights (as well as other environmental information such as weather conditions, road surface conditions, etc.).
[0034] Perform data preprocessing and integration on the obtained environmental observation information to construct a state space and an action space, providing input for the multi-agent reinforcement learning algorithm. Among them, the process of data preprocessing and integration is as follows:
[0035] Step 1.1. Construction of the state space and definition of the action space.
[0036] Construction of the state space:
[0037] Integrate the obtained environmental observation information into a state vector as the observation input for each AVs. The state vector includes the state of the current vehicle, the state of observable surrounding vehicles, and environmental information.
[0038] The state of the current vehicle includes longitudinal position, lateral position, longitudinal speed, and lateral speed. The state of observable surrounding vehicles includes relative longitudinal position, relative lateral position, relative longitudinal speed, and relative lateral speed. Environmental information includes road topological structure, traffic signs, and traffic lights. Define all execution actions of the vehicle AVs, including lateral control and longitudinal control; lateral control includes keeping in the lane, changing lanes to the left, and changing lanes to the right; longitudinal control includes maintaining speed, accelerating, and decelerating.
[0039] Step 1.2. Data Storage and Management.
[0040] Use an experience replay buffer to store recent experimental data, including states, actions, rewards, and value estimates; the size of the experience replay buffer is set to 100, and only the most recent 100 steps of data are stored for model training.
[0041] Step 1.3. Data Usage and Parameter Update.
[0042] Data Usage during Training:
[0043] During each training episode, the current vehicle selects an action according to the current policy, and after executing the action, obtains new state, reward, and other information, which are stored in the experience replay buffer.
[0044] Parameter Update and Data Clearing:
[0045] When the episode ends or a collision occurs, sample a mini-batch of data from the buffer to calculate the gradients of the policy network and the value network, and update the parameters; then clear the experience replay buffer to prepare for data collection in the next training episode.
[0046] Step 2. Design a safety awareness module based on safety-enhanced multi-agent reinforcement learning, as Figure 2 shown, which improves the safety of vehicles in the high-speed merging area. Among them, the design process of the safety awareness module is as follows:
[0047] Construct a probability model to model the Time To Collision (TTC) and the Hidden Zone Threat (HZT) respectively, and use a piecewise likelihood function combined with Gaussian and uniform distributions to flexibly capture environmental uncertainties.
[0048] TTC comprehensively considers the time interval and relative speed between vehicles, providing an accurate assessment of potential collision risks, while HZT focuses on detecting and evaluating potential threats within the vehicle's blind spot, enhancing the system's ability to perceive invisible risks.
[0049] The present invention combines these two indicators, TTC and HZT, to improve the comprehensiveness and reliability of safety analysis, ensuring that the system can quickly perceive risks and optimize decisions in a complex dynamic environment.
[0050] The process of modeling the collision time TTC and the blind spot threat HZT is as follows:
[0051] The calculation formula for the key indicator collision time TTC is:
[0052] .
[0053] Among them, represents the time to collision (TTC) of the current vehicle. , denotes the set of AVs; , o represents the observed vehicle, m represents the vehicle on the main road closest to the current vehicle, and r represents the vehicle on the ramp closest to the current vehicle; and represent the longitudinal distance and safety distance between the current vehicle and vehicle o, and represent the speeds of the current vehicle and the observed vehicle o, respectively. Define as the time to collision of the closest vehicle on the main road and the time to head-on collision of the closest vehicle on the ramp as the minimum value between them, and its formula is:
[0054] .
[0055] The calculation formula for the hidden zone threat (HZT) of the blind area is:
[0056] .
[0057] Among them, represents the hidden zone threat in front of the current vehicle, represents the lateral position deviation between the current vehicle and the observed vehicle o, represents the lateral blind area threshold of the current vehicle.
[0058] Finally is defined as the minimum value between the hidden zone threat in front of the closest vehicle on the main road and the hidden zone threat of the closest vehicle on the ramp : .
[0059] To evaluate the safety level of AVs, it is necessary to set the thresholds of the safety awareness module. The present invention sets the safety thresholds for TTC and HZT. These thresholds are determined based on historical data and expert knowledge and are used to quantify the safety state of the vehicle. Specifically, the TTC threshold is calculated according to the average reaction time and braking distance of the vehicle. When TTC is less than this threshold, it is considered that the vehicle has enough time to avoid a collision and is in a safe state. The HZT threshold is determined according to the size of the vehicle's blind area and the distribution of observable vehicles around. When HZT is lower than this set threshold, it indicates that there is no potential threat in the vehicle's blind area and the vehicle is in a safe state. When and exceed the threshold, the vehicle is usually in a safe state but may still face certain risks. On the contrary, when and When below the threshold, due to insufficient braking time, short braking distance, or obstructed visibility, the vehicle may face a collision risk.
[0060] To effectively characterize the uncertainty of the collision risk, the present invention uses a probability density function for modeling and a Gaussian distribution for reasonable approximation. Therefore, a piecewise likelihood function is used, which consists of a uniform distribution within a defined interval and two Gaussian components outside this interval. This method can flexibly represent environmental uncertainty and allows for continuous learning and dynamic adjustment of the safety level assessment in the absence of prior information or in case of unexpected situations.
[0061] Using the above TTC threshold and HZT threshold, the present invention defines three safety levels:
[0062] Safety level : When both TTC and HZT are far higher than the corresponding thresholds, the vehicle is in the safest state. Medium level : When TTC or HZT approaches the threshold, the vehicle is in a medium safety state and needs to drive carefully. Dangerous level : When both TTC or HZT are below the corresponding thresholds, the vehicle is in a dangerous state and immediate measures need to be taken to avoid a collision.
[0063] For the safety levels of safe , medium , dangerous , different forms of likelihood functions are constructed respectively, and the probability distribution of the vehicle safety level is updated through Bayes' formula. Finally, the vehicle safety level is determined by maximizing the posterior probability. Specifically, the process of constructing different forms of likelihood functions for the safety levels of safe , medium , dangerous is as follows:
[0064] The formula for the likelihood function of the dangerous level is as follows:
[0065] .
[0066] Where represents the forward threat metric at the collision risk assessment level . represents the uncertainty of vehicle k, represents the uncertainty of vehicle k, represents the collision time threshold of the leading vehicle at the dangerous safety level , Indicates the blind spot threat threshold of the vehicle ahead at the dangerous safety level The first Gaussian distribution is used to quantify the risk of TTC deviating from the safety threshold, and the second Gaussian distribution is used to quantify the risk of HZT deviating from the safety threshold, with equal weights
[0067] indicating that TTC and HZT contribute equally to the danger level
[0068] For the medium level The formula for the likelihood function is as follows
[0069]
[0070] Where Indicates the forward threat metric at the collision risk assessment level Indicates the time-to-collision threshold of the vehicle ahead at the medium level Indicates the blind spot threat threshold of the vehicle ahead at the medium level
[0071] The medium risk state needs to consider both gradual and sudden factors. Therefore, regardless of whether and exceed the threshold, the same form of Gaussian distribution is used, but the variance is larger, indicating a higher tolerance for deviation
[0072] For the safe level The formula for the likelihood function is as follows
[0073]
[0074] Where Indicates the forward threat metric at the collision risk assessment level Indicates the time-to-collision threshold of the vehicle ahead at the safe level Indicates the blind spot threat threshold of the vehicle ahead at the safe level
[0075] In the safe state, only when both TTC and HZT are far higher than the threshold, a loose Gaussian distribution is used, that is, the variance is the largest, otherwise it is directly assigned a value of 1, indicating no risk. Where Indicates the safety assessment measure is an element of the safety assessment measure set denote threshold value denote threshold value
[0076] For vehicle k and the uncertainties are denoted by and respectively
[0077] The piecewise design avoids the integration operation of complex mixed distributions, and the posterior probability is updated quickly through Bayes' formula. The probability distribution of the vehicle safety level is dynamically updated according to the observed data:
[0078] .
[0079] where j ∈ {0, 1, 2}, and j represents the index of the vehicle safety level denotes the posterior probability that the vehicle is at the safety level index j, while denotes the prior probability that the vehicle is at the safety assessment level j, the prior probability that the vehicle is at the safety assessment level denotes the safety level of the vehicle ahead under the safety level index j denotes the collision time of the vehicle ahead under the safety level index j denotes the blind spot threat of the vehicle ahead under the safety level index j denotes the i-th safety level, where i ranges from 1 to 3, representing different safety levels
[0080] Finally, the safety level is determined by maximizing the posterior probability:
[0081] .
[0082] where denotes the safety level of the vehicle ahead determined by maximizing the posterior probability denotes selecting the safety level that maximizes the posterior probability denotes the set composed of all possible safety levels
[0083] Step 3. Design a multi-agent decision-making module, as shown in Figure 3 to achieve multi-vehicle cooperation
[0084] Define the state space, action space, and reward function. The state vector of each AV is used to comprehensively describe the surrounding environment. The action space includes six discrete actions. The reward function includes safety rewards, collision rewards, speed rewards, time headway rewards, and lane change rewards. Construct a shared network structure, including an actor network and a critic network, which share the same low-level representation, including position and speed information, to facilitate the optimization of the policy gradient update and value function learning process by the SMA2C algorithm.
[0085] Step 3 specifically includes:
[0086] Step 3.1. The state space S of each AV i is defined as a matrix with dimension N i × s i where N i represents the number of observable vehicles, and the state vector s i represents the state vector of the vehicle at the current time step; s i includes the longitudinal position x, lateral position y, longitudinal speed v x , and lateral speed v y with respect to the current vehicle. Define s i = {x, y, v x , v y}.
[0087] Assume that the number of observable vehicles N i of the current vehicle is finite and is defined as the number of the nearest vehicles within a longitudinal distance of 150 meters from the current vehicle; in the lane change scenario, the best performance can be achieved when N i = 5.
[0088] The joint state space is defined as: .
[0089] The action space a of each AV i = {Keep Lane, Lane Change Left, Lane Change Right, Keep Speed, Accelerate, Decelerate}.
[0090] The joint state space is defined as .
[0091] where the subscript i is used to distinguish different AVs and corresponds to their state vectors s i and action spaces a i respectively.
[0092] Step 3.2. Design safety rewards, collision rewards, speed rewards, time headway rewards, and lane change rewards.
[0093] The reward function is a dynamic reward function designed based on the safety level output by the safety awareness module. It dynamically adjusts the reward value to address the dynamic interaction problem in the highway merge area, while other reward functions are static.
[0094] The present invention adopts the combination of a dynamic reward function and a static reward function. Among them, the static reward can provide the advantages of stability and clear goals, while the dynamic reward can flexibly respond to changes in the environment and the current vehicle state.
[0095] At time step t, the reward of the autonomous vehicle i is defined as follows:
[0096] .
[0097] Where represents the safety reward, represents the collision reward, represents the speed reward, represents the time headway reward, represents the lane change reward, , , , , are the weighting coefficients of the corresponding rewards.
[0098] In this embodiment, the weighting coefficients of the reward function are set to .
[0099] The safety reward is used to evaluate whether the current vehicle is performing tasks while ensuring safety and avoiding collisions:
[0100] ;
[0101] Where represents the forward safety level, represents the reverse safety level; and respectively represent the weights of the forward safety level and the backward safety level ; because the forward and backward safety levels are equally important, = .
[0102] In this embodiment, is calculated in the same way as .
[0103] The collision reward , which means that a collision will be punished, and its formula is as follows:
[0104] 。
[0105] Speed Reward , which is used to evaluate whether the current vehicle maintains an appropriate speed range. The formula is as follows:
[0106] 。
[0107] Wherein, is the current speed of the current vehicle. The minimum and maximum speeds of the current vehicle are set to and ; In this embodiment, the minimum and maximum speeds of the current vehicle are set to = 10 m / s, = 30 m / s.
[0108] Time Headway Reward , which is used to ensure that the current vehicle maintains a safe distance from the vehicle in front and avoid potential collision risks caused by insufficient headway. Its calculation formula is as follows:
[0109] ;
[0110] Wherein, is the distance headway, is the predefined time headway threshold; In this embodiment = 1.2 s.
[0111] Lane Change Reward , which is used to encourage the current vehicle to use the merging lane and improve efficiency through operations:
[0112] 。
[0113] Wherein, is the distance that the current vehicle travels on the ramp, is the length of the ramp. As the current vehicle approaches the merging end of the highway on-ramp, the penalty (when x approaches the ramp length L, that is, when the vehicle approaches the merging point of the highway main road, the value of the exponential function decreases rapidly, resulting in becoming more negative, that is, the penalty increases) also increases.
[0114] Step 3.3. Construct an actor network and a critic network, sharing the same low-level representation, including position state and speed state. These states are input into separate fully connected layers for processing and then combined, and then passed to the final fully connected layer to output the action probability distribution for the actor network and the state value for the critic network.
[0115] The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning optimizes the policy gradient update and value function learning process by introducing a security perception module and a multi-agent cooperation mechanism on the basis of the traditional A2C (Advantage Actor-Critic) framework.
[0116] Step 3.3.1. Policy gradient calculation, combining the output of the security perception module, i.e., the security level and , dynamically adjusts the weight of the advantage function, increasing the preference for safe actions in high-risk states:
[0117] .
[0118] Among them represents the policy gradient; |B| represents the batch size, i.e., the number of samples extracted from the experience buffer; represents the probability of selecting action in state at time step t; represents the action value function, represents the baseline value, with t as the subscript to distinguish different moments of the current agent.
[0119] Step 3.3.2. Relationship between the action value function and the state value function, comes from the critic network; however, due to the potential estimation bias introduced by the double-network structure, the baseline is selected as the value function .
[0120] To further reduce the estimation variance, the action value function is also represented by the value function, and the formula is as follows:
[0121] .
[0122] Among them, represents the reward obtained at time step t; represents the state value function of state , representing the expected cumulative return obtained when starting from state and following the current policy π.
[0123] Step 3.3.3. Advantage function , used to measure the gain of taking action in state relative to the baseline value function , and can be used to update the parameters of the critic network based on the time difference error, and the formula is as follows:
[0124] .
[0125] Step 3.3.4. Policy Gradient Update Formula and parameters update rules, which aim to optimize the policy by adjusting the parameters in the desired direction. The calculation formulas are shown as follows:
[0126] ; .
[0127] Among them, represents the learning rate, which controls the step size for updating the policy parameters .
[0128] Step 3.3.5. Value Function Loss , which is defined as the mean squared error of the time difference error. The calculation formula is as follows
[0129] .
[0130] Among them, |B| represents the batch size, that is, the number of samples extracted from the experience buffer; represents the reward obtained at time step t; , respectively represent the state , state 's state value function; represents the discount factor.
[0131] Step 3.3.6. Policy Gradient of the Value Network The calculation formula and parameters update rules are as follows:
[0132] ;
[0133] .
[0134] Among them, represents the learning rate, which controls the step size for updating the value network parameters .
[0135] Step 4. Introduce an action masking module, as shown in Figure 4 , which is used to mask the unsafe actions in the output of the multi-agent decision-making module to ensure that the system only selects safe operations. Action masking is a technique that restricts the selection of invalid or unsafe actions in reinforcement learning (RL) by using a mask, ensuring that the agent only selects valid actions.
[0136] This method avoids the selection of invalid actions and only samples from valid actions.
[0137] As shown in Figure 4As shown, the action masking module is implemented through the following steps:
[0138] Step 4.1. Define an action masking vector for masking unsafe actions, including:
[0139] Transform AVs to non-existent lanes;
[0140] When the speed of AVs has reached the preset maximum or minimum speed, further accelerate or decelerate; by predefined this vector, restrict the selection of unsafe actions to ensure that the system only selects safe and effective actions.
[0141] In the action masking vector, 0 represents an invalid or unsafe action, and 1 represents a valid action.
[0142] Step 4.2. Apply action masking:
[0143] In the output of the decision-making module, set the logits of unsafe actions to very small negative values so that after these actions are normalized by the Softmax function, the selected probability is close to 0.
[0144] In addition, to verify the effectiveness of the method proposed in the present invention, the present invention also gives the following experimental environment scenario design. To evaluate the performance of the vehicle under different traffic densities, the scenario is set as follows:
[0145] I. At low traffic density, quickly identify the merging opportunity without affecting the vehicles on the main road;
[0146] II. At medium traffic density, find gaps among multiple vehicles and cope with speed changes and lane changes;
[0147] III. At high traffic density, safely complete the merge within a limited space and time to avoid congestion or collision;
[0148] IV. When there are multiple AVs, make effective collaborative decisions to avoid conflicts or efficiency reduction.
[0149] Specifically, in this embodiment, the simulation environment is modified based on the Gym's highway-env simulator.
[0150] The scenario includes the main road and the ramp. The total length of the main road is 520 meters, of which the passing sections (AB and DE sections) are 320 meters long, the ramp entrance section (BC section) is 100 meters long, and the ramp merging area (CD section) is 100 meters long. The starting points of the main road and the ramp are the same, and the vehicle generation point is set on the passing lane of the AB section. Vehicles outside the range will be removed.
[0151] In this embodiment, three traffic density scenarios are designed, namely: low traffic density: 1 - 3 AVs and 1 - 3 HDVs; medium traffic density: 2 - 4 AVs and 2 - 4 HDVs; high traffic density: 4 - 6 AVs and 3 - 5 HDVs.
[0152] A specific traffic scenario is set in the highway - env simulator to evaluate the performance of vehicles under the following scenarios:
[0153] I. In the case of low traffic density, the vehicle needs to quickly identify and utilize appropriate opportunities to merge in, while avoiding affecting the normal driving of the main - road vehicles.
[0154] II. In the case of medium traffic density, the vehicle needs to find appropriate gaps among multiple HDVs to merge in, while coping with the speed changes and lane - changes of other vehicles.
[0155] III. In the case of high traffic density, the vehicle needs to safely complete the merging operation within limited space and time, avoiding causing traffic congestion or collision accidents.
[0156] IV. In the scenario with multiple AVs, effective collaborative decision - making is required among vehicles to achieve efficient and safe merging, avoiding conflicts or reduced efficiency caused by multiple AVs competing for the same merging gap simultaneously.
[0157] The method of the present invention is based on deep reinforcement learning. By introducing a safety - awareness module and a multi - agent decision - making module, the safety and merging efficiency of AVs in the highway merging area are significantly improved. Through parallelization and distributed processing, multiple AVs can simultaneously execute policy training in different environments, sharing policy updates and experience data in real - time, thereby accelerating the convergence speed of the algorithm and reducing the time cost required for a single AV to conduct independent training. In addition, the method of the present invention also combines action - masking technology to further enhance the diversity and effectiveness of exploration, avoiding premature convergence to sub - optimal policies. In the multi - agent architecture, collaborative work among AVs is realized, ensuring that the exploration process is more comprehensive on a global scale. Combining the design of the safety - awareness module and the multi - agent decision - making module, the method of the present invention significantly improves the efficiency of the training process, enhances the robustness and generalization ability of the policy, and makes it perform better in complex autonomous - driving decision - making tasks. The method of the present invention can effectively cope with the changes in the highway merging - area environment and maintain good merging efficiency and safety.
[0158] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above - mentioned embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.
Claims
1. An autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning, characterized in that It includes the following steps: Step 1. Obtain environmental observation information including the current vehicle state information, observable surrounding vehicle information, and environmental information, perform data preprocessing and integration on the obtained environmental observation information, and construct a state space and an action space; Step 2. Design a safety awareness module based on safety-enhanced multi-agent reinforcement learning; First, construct a probability model, model the time to collision (TTC) and the blind spot threat (HZT) respectively, and adopt a piecewise likelihood function combined with Gaussian and uniform distributions to flexibly capture environmental uncertainties; For different safety levels, construct different forms of likelihood functions respectively, update the probability distribution of the vehicle safety level through Bayes' formula, and finally determine the vehicle's safety level by maximizing the posterior probability; Step 3. Design a multi-agent decision-making module to achieve multi-vehicle cooperation; Define the state space, action space, and reward function; The state vector of each AVs is used to comprehensively describe the surrounding environment. The action space includes 6 discrete actions; the reward function includes safety rewards, collision rewards, speed rewards, time headway rewards, and lane change rewards; Construct a shared policy network structure, including an actor network and a critic network, which share the same low-level representation, including position state and speed state, to optimize the policy gradient update and value function learning process; Step 4. Introduce an action masking module, and through predefined rules, mask invalid or unsafe actions to limit the output of the multi-agent decision-making module to ensure that the system performs safe operations.
2. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 1, wherein In the above Step 1, the process of obtaining environmental observation information is as follows: I. Obtaining vehicle state information: For each autonomous vehicle AVs, obtain its state information, including position, speed, and acceleration; II. Obtaining observable surrounding vehicle information: Obtain the position and speed information of human-driven vehicles (HDVs) and AVs within the observable range around each AVs to construct the observation state of the vehicle; III. Obtaining environmental information: Obtain environmental information of the road, including topological structure, traffic signs, and traffic lights.
3. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 1, characterized in that, In the above Step 1, the process of performing data preprocessing and integration on the obtained environmental observation information is as follows: Step 1.
1. Constructing the state space and defining the action space; Integrate the obtained environmental observation information into a state vector as the observation input of each AVs. The state vector includes the state of the current vehicle, the state of observable surrounding vehicles, and environmental information; The state of the current vehicle includes longitudinal position, lateral position, longitudinal speed, and lateral speed; the state of observable surrounding vehicles includes relative longitudinal position, relative lateral position, relative longitudinal speed, and relative lateral speed; Environmental information includes road topological structure, traffic signs, and traffic lights; Define all execution actions of each AVs, including lateral control and longitudinal control; lateral control includes keeping lane, changing lane to the left, and changing lane to the right; longitudinal control includes keeping speed, accelerating, and decelerating; Step 1.
2. Data storage and management; Use an experience replay buffer to store recent experimental data, including states, actions, rewards, and value estimates; the size of the experience replay buffer is set to 100, and only the most recent 100 steps of data are stored for model training; Step 1.
3. Data usage and parameter update; In each training episode, the current vehicle selects an action according to the current policy, and after executing the action, obtains new state, reward and other information, which are stored in the experience replay buffer; When the episode ends or a collision occurs, a mini-batch of data is sampled from the buffer to calculate the gradients of the policy network and the value network, and the parameters are updated; subsequently, the experience replay buffer is cleared to prepare for data collection in the next training episode.
4. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 1, wherein, In Step 2, the process of modeling the time-to-collision TTC and the blind-zone threat HZT is as follows: The calculation formula for the key metric time-to-collision TTC is: ; Among them, represents the time to forward collision of the current vehicle; , represents the set of AVs; , o represents the observed vehicle, m represents the vehicle on the main road closest to the current vehicle, and r represents the vehicle on the ramp closest to the current vehicle; , represent the longitudinal distance and safety distance between the current vehicle and vehicle o, , represent the speeds of the current vehicle and the observed vehicle o respectively; define as the time to forward collision of the closest vehicle on the main road and the time to frontal collision of the closest vehicle on the ramp the minimum value between them, and its formula is: ; The calculation formula for the blind-zone threat HZT is: ; Among them, represents the threat in the front hidden area of the current vehicle, indicates the lateral position deviation between the current vehicle and the observed vehicle o, represents the lateral blind area threshold of the current vehicle; Final is defined as the threat of the front blind spot of the nearest vehicle on the main road and the threat of the blind spot of the nearest vehicle on the ramp The minimum value between: .
5. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 1, characterized in that In Step 2, the process of constructing different likelihood function forms for three different safety levels is as follows: Danger level The formula for the likelihood function is as follows: ; Among them, represents the forward threat metric at the collision risk assessment level; ; represents the uncertainty of vehicle k, represents the uncertainty of vehicle k, represents the collision time threshold of the leading vehicle at the hazard safety level, represents the blind spot threat threshold of the leading vehicle at the hazard safety level; The first Gaussian distribution is used to quantify the risk of TTC deviating from the safety threshold, and the second Gaussian distribution is used to quantify the risk of HZT deviating from the safety threshold, with equal weights of indicating that TTC and HZT contribute equally to the risk level; Medium level The formula for the likelihood function is as follows: ; Among them, represents the forward threat metric under the collision risk assessment level ; represents the time-to-collision threshold when the vehicle ahead is at the medium level represents the blind spot threat threshold when the vehicle ahead is at the medium level The medium-risk state needs to consider both gradual and sudden factors. Therefore, regardless of and whether it exceeds the threshold, the same form of Gaussian distribution is adopted, but the variance is larger, indicating a higher tolerance for deviation; Safety level The formula for the likelihood function is as follows: ; Among them, represents the forward threat metric at the collision risk assessment level ; represents the time-to-collision threshold of the forward vehicle at the safety level ; represents the blind spot threat threshold of the forward vehicle at the safety level In the safe state, only when both TTC and HZT are higher than their corresponding thresholds, a loose Gaussian distribution with the largest variance is adopted; otherwise, it is directly assigned a value of 1, indicating no risk. Among them, represents the safety assessment measure, which is an element of the safety assessment measure set . represents the threshold of , and represents the threshold of . is the largest, otherwise it is directly assigned a value of 1, indicating no risk; among them represents the safety assessment measure, is a set of safety assessment measures is an element of represents the threshold of represents the threshold; Of vehicle k and The uncertainties of are respectively represented by and ; Piecewise design avoids the integral operation of complex mixed distributions, and the posterior probability is updated quickly through Bayes' formula, and the probability distribution of the vehicle safety level is dynamically updated according to the observed data: ; where \(j\in\{0,1,2\}\), and \(j\) represents the index of the safety level of the vehicle; represents the posterior probability that the vehicle is at the safety level index \(j\), while represents the prior probability that the vehicle is at the safety assessment level \(j\), and the prior probability that the vehicle is at the safety assessment level; represents the safety level of the vehicle ahead under the safety level index \(j\), represents the collision time of the vehicle ahead under the safety level index \(j\), represents the blind spot threat of the vehicle ahead under the safety level index \(j\), represents the \(i\)-th safety level, where \(i\) ranges from 1 to 3, representing different safety levels; Finally, the safety level is determined by maximizing the posterior probability: ; Among them, represents the safety level of the vehicle ahead determined by maximizing the posterior probability, represents selecting the safety level that maximizes the posterior probability, represents the set composed of all possible safety levels.
6. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 1, wherein Step 3 includes the following steps: Step 3.
1. The state space S of each AVs i is defined as a matrix of dimension N i × s i ; where N i represents the number of observable vehicles, and the state vector s i represents the state vector of the vehicle at the current time step; s i includes the longitudinal position x, lateral position y, longitudinal velocity v x with respect to the current vehicle, and lateral velocity v y ; The combined state space is defined as: ; The action space \(a\) of each AV i = {Keep Lane, Lane Change Left, Lane Change Right, Keep Speed, Accelerate, Decelerate}; The combined state space is defined as ; where the subscript \(i\) is used to distinguish different AVs, corresponding to their state vectors \(\mathbf{s}\) i and action space \(\mathbf{a}\) i ; Step 3.
2. Design safety rewards, collision rewards, speed rewards, time headway rewards, and lane-changing rewards; The safety reward is designed as a dynamic reward function based on the safety level output by the safety awareness module, and the other rewards in the reward function are static reward functions. The combination of dynamic and static reward functions provides stability and flexibility; Step 3.
3. Construct an actor network and a critic network, which share the same low-level representation, including position state and speed state. These states are input into separate fully connected layers for processing and then combined, and then passed to the final fully connected layer to output the action probability distribution for the actor network and the state value for the critic network.
7. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 6, characterized in that, In step 3.2, at time step t, the reward of autonomous vehicle i is defined as follows: ; Among them, represents the safety reward, represents the collision reward, represents the speed reward, represents the time headway reward in the workshop, represents the lane change reward, , , , , are the weighting coefficients of the corresponding rewards; Safety Reward For evaluating whether the current vehicle performs tasks while ensuring safety and avoiding collisions: ; Among them, represents the forward safety level, and represents the reverse safety level; and represent the weights of the forward safety level and the backward safety level respectively; since the forward and backward safety levels are equally important, = ; Collision Reward , which means that a penalty will be imposed for collisions, and the formula is as follows: ; Speed Reward , which is used to evaluate whether the current vehicle is maintaining an appropriate speed range. The formula is as follows: ; Among them, is the current speed of the current vehicle, and the minimum and maximum speeds of the current vehicle are set to and ; Workshop Time Headway Reward , which is used to ensure that the current vehicle maintains a safe distance from the vehicle in front and avoid potential collision risks caused by insufficient time headway. Its calculation formula is as follows: ; Among them, is the time headway from the vehicle head, is the predefined time headway threshold value; Lane change reward , which is used to encourage the current vehicle to use the merging lane and improve efficiency through operations: ; Among them, is the distance that the current vehicle travels on the ramp, is the length of the ramp.
8. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 6, characterized in that, The specific design process of Step 3.3 is as follows: Step 3.3.
1. Policy gradient calculation, combining the output of the safety perception module, i.e., the safety level and , dynamically adjust the weight of the advantage function, increasing the preference for safe actions in high-risk states: ; Among them represents the policy gradient; |B| represents the batch size, i.e., the number of samples extracted from the experience buffer; represents the probability of selecting action at state at time step t; represents the action value function, and represents the baseline value; t is used as a subscript to distinguish different moments of the current vehicle; Step 3.3.
2. Relationship between the action value function and the state value function comes from the critic network; however, due to the potential estimation bias introduced by the dual network structure, the baseline is selected as the value function ; To further reduce the estimation variance, the action value function is also represented by the value function, and the formula is as follows: ; where, represents the reward obtained at time step t; represents the state state value function of, representing the expected cumulative return obtained when starting to follow the current policy π from state ; Step 3.3.
3. Advantage function , used to measure the state Relative to baseline value function Take Action The gain can be used to update the parameters of the critic network based on the time difference error, as follows: ; Step 3.3.
4. Policy Gradient Update Formula and parameters update rules, which aim to optimize the policy by adjusting the parameters in the desired direction. The calculation formulas are shown as follows respectively: ; ; Among them, represents the learning rate, which controls the step size for updating the policy parameters ; Step 3.3.
5. Value function loss , which is defined as the mean squared error of the temporal difference error, and the calculation formula is as follows ; Among them, |B| represents the batch size, that is, the number of samples extracted from the experience buffer; represents the reward obtained at time step t; , respectively represent the state , state value function; represents the discount factor; Step 3.3.
6. Policy Gradient of Value Network Calculation formula and parameters The update rule is as follows: ; ; Among them, represents the learning rate, which controls the step size for updating the parameters of the value network of the step.
9. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 1, characterized in that Step 4 includes the following steps: Step 4.
1. Define an action masking vector to mask unsafe actions, including: Transforming AVs to non-existent lanes; When the speed of AVs has reached the preset maximum or minimum speed, further accelerating or decelerating; In the action masking vector, 0 represents an invalid or unsafe action, and 1 represents a valid action; Step 4.
2. Apply action masking: In the output of the multi-agent decision-making module, set the logarithm of the unsafe actions to a small negative value, so that after these actions are normalized by the Softmax function, the probability of being selected is close to 0.
10. The autonomous driving decision-making method based on security-enhanced multi-agent reinforcement learning according to claim 1, wherein After Step 4, it also includes: Step 5. Design an experimental environment scenario to evaluate the performance of the vehicle under different traffic densities. The scenario settings are as follows: I. At low traffic density, quickly identify the merging opportunity without affecting the vehicles on the main road; II. At medium traffic density, find gaps among multiple vehicles and cope with speed changes and lane changes; III. At high traffic density, safely complete the merge within limited space and time to avoid congestion or collisions; IV. Make effective collaborative decisions when there are multiple AVs to avoid conflicts or reduced efficiency.
Citation Information
Patent Citations
Automatic driving vehicle risk assessment and interactive planning system in perception blind area scene
CN118536793A
Intelligent highway lane changing method for autonomous vehicle based on reinforcement learning
CN119568155A