Autonomous driving decision-making method, system and vehicle with dynamic risk perception and attention focus
By employing a dual-deep Q-network enhanced with a multi-feature Gaussian weighted particle filter algorithm and an attention mechanism, the risk perception and decision-making problems of autonomous driving systems in complex traffic environments are solved, thereby improving the safety and efficiency of decision-making.
Patent Information
- Application Number
- CN202511141594.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing autonomous driving decision-making systems struggle to achieve accurate risk perception and intelligent decision-making in complex and ever-changing traffic environments, resulting in low safety and efficiency.
A multi-feature Gaussian weighted particle filter algorithm is used for trajectory prediction. Combined with the vehicle kinematics model and multi-dimensional error calculation, a comprehensive driving risk assessment model is designed, and a dual-deep Q-network with attention mechanism enhancement is introduced for decision optimization.
It improves the safety and reliability of autonomous driving systems in complex traffic environments, enabling accurate identification and efficient decision-making of potential hazards.
Smart Images

Figure CN120716726B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of automatic driving, and particularly relates to an automatic driving decision method and system with dynamic risk perception and attention focusing, and a vehicle. BACKGROUND
[0002] In recent years, automatic driving technology has made significant progress. Various perception, decision and control systems have been continuously improved, enabling autonomous driving vehicles to achieve autonomous driving in different scenarios. Among them, the decision system is one of the core modules of the autonomous driving vehicle, which is responsible for planning the future trajectory and action of the vehicle according to the perceived environmental information.
[0003] Existing automatic driving decision systems are mainly designed based on rules or models. Rule-based decision systems guide the driving behavior of vehicles through a pre-set rule base, for example, when a front vehicle is detected to slow down, a slow-down operation is performed. Model-based decision systems use machine learning and other technologies to train models to learn the driving behavior of drivers, and make decisions based on the predicted results of the model.
[0004] However, the existing automatic driving decision system still has some defects, which limits its application range and safety. First, rule-based decision systems are difficult to cope with complex and variable traffic environments, for example, when unexpected situations occur, the rule base may not provide effective decisions. Second, model-based decision systems often require a large amount of training data, and the generalization ability of the model is limited, making it difficult to adapt to different driving scenarios. In addition, the existing decision system usually lacks dynamic risk perception ability, and cannot adjust the decision strategy in time according to the current environmental state and traffic situation, which may lead to safety accidents. SUMMARY
[0005] The purpose of the present application is to provide an automatic driving decision method and system with dynamic risk perception and attention focusing, to realize accurate risk perception and intelligent decision-making ability of autonomous driving vehicles in complex dynamic traffic environments, to overcome the defects of insufficient quantification of uncertain factors, incomplete risk assessment and low decision-making efficiency of existing methods.
[0006] The present application achieves the above-mentioned purpose through the following technical solutions:
[0007] The present application provides an automatic driving decision method with dynamic risk perception and attention focusing, comprising the following steps:
[0008] S1. initializing a particle set based on a current observed vehicle state, updating a particle set state through a vehicle kinematics model and introducing Gaussian noise to simulate uncertainty, calculating particle weights by using a multi-dimensional Gaussian distribution, and outputting a weighted average predicted trajectory after optimization by resampling; the particle weights include position error, heading angle error and speed error;
[0009] S2. calculating a longitudinal collision risk value and a lateral collision risk value based on a current position, a predicted position, a speed and a safety distance of the vehicle, and dynamically adjusting weights according to traffic density to fuse and output an integrated driving risk value IDR;
[0010] S3. inputting the integrated driving risk value IDR and a vehicle driving intention after feature integration into a double deep Q network based on an attention mechanism enhancement to generate a lane changing decision of the autonomous vehicle.
[0011] As a preferred scheme of the present application, step S1 comprises:
[0012] S11. acquiring real-time state data of the observed vehicle at the current time, including vehicle position coordinates, a heading angle and a speed;
[0013] S12. generating N initial particles in a three-dimensional state space according to the current observed state data, and each particle state obeys a Gaussian distribution with the observed value as a mean value;
[0014] S13. for each particle, calculating a next time state prediction value according to a vehicle kinematics model;
[0015] S14. for each predicted particle, calculating three-dimensional errors thereof with the latest observed data, including position error, heading angle error and speed error;
[0016] S15. calculating particle weights based on the three-dimensional errors, performing normalization processing on the weights, and performing system resampling to reserve high-weight particles;
[0017] S16. outputting a weighted average state of the resampled particles as a final prediction result.
[0018] As a preferred scheme of the present application, in step S13, the particle state update adopts the following kinematics model:
[0019]
[0020] wherein are a predicted position, a predicted heading angle and a predicted speed of the vehicle respectively; are a current position, a current heading angle and a current speed of the vehicle respectively; is a yaw angle, is a prediction time step, L is a wheelbase of the vehicle, is a Gaussian noise, are variances of the current acceleration and the Gaussian noise of the ego vehicle, respectively.
[0021] As a preferred scheme of the present application, the step S2 comprises:
[0022] S21. obtaining current states and predicted states of the ego vehicle and the surrounding vehicle;
[0023] S22. calculating a longitudinal risk value as follows:
[0024]
[0025] wherein and represent longitudinal positions of the surrounding vehicle and the ego vehicle, respectively, and represent predicted longitudinal positions of the surrounding vehicle and the ego vehicle, respectively; is a longitudinal safety distance between the ego vehicle and the surrounding vehicle, is a longitudinal real distance between the ego vehicle and the surrounding vehicle; is a maximum value of the longitudinal risk, is a shape adjustment factor; are dynamic convergence coefficients and predicted dynamic convergence coefficients along a longitudinal direction, respectively, wherein is a reference speed, is a current speed of the ego vehicle, is a predicted speed of the ego vehicle, is a constant coefficient, is a longitudinal constant convergence coefficient;
[0026] S23. calculating a lateral risk value as follows:
[0027]
[0028] wherein is a maximum value of the lateral risk, is a lateral safety distance between the ego vehicle and the surrounding vehicle, which is set based on a road width; a lateral real distance wherein is a lateral distance between the ego vehicle and the surrounding vehicle, represents a distance from a center of mass of the ego vehicle to a side of the vehicle body; and are dynamic convergence coefficients and predicted dynamic convergence coefficients along a lateral direction, respectively, wherein is a constant coefficient, is a lateral constant convergence coefficient;
[0029] S24. Calculate the comprehensive risk value IDR as follows:
[0030]
[0031] wherein and are the dynamic weight coefficients of the longitudinal and lateral risks related to the traffic density, , ; wherein is the current traffic density, which is calculated according to the ratio of the number of vehicles to the number of roads within a certain range. and are the minimum and maximum values of the traffic density, respectively.
[0032] As a preferred scheme of the present application, in step S22, the calculation of the longitudinal safety distance and the real distance comprises the following formula:
[0033]
[0034] wherein and represent the speed and predicted speed of the ego vehicle and surrounding vehicles, respectively; and represent the heading angle and predicted heading angle of the ego vehicle and surrounding vehicles, respectively, is the longitudinal distance between the ego vehicle and the surrounding vehicle, is the distance from the center of mass of the ego vehicle to the front of the vehicle, represents the distance from the center of mass of the ego vehicle to the side of the vehicle, is the safety distance margin, is the driver reaction time, is the vehicle braking system reaction time, is the maximum deceleration of the vehicle.
[0035] As a preferred scheme of the present application, step S3 comprises:
[0036] S31. Encode the ego vehicle and surrounding vehicle states into feature vectors;
[0037] S32. Calculate the similarity weight between the ego vehicle and the surrounding vehicles through the self-attention mechanism, and generate the environment features representing the attention distribution by weighted fusion;
[0038] S33. Concatenate the features with the original ego vehicle state, input the DDQN network to generate the action policy, which includes left lane change, right lane change, acceleration, deceleration and keep.
[0039] As a preferred scheme of the present application, the method further comprises reward function design, specifically comprising:
[0040] A safety reward function is set to punish the speed of the ego vehicle positively when a collision occurs, as follows:
[0041]
[0042] is a safety reward function, where is the forward speed of the ego vehicle, is the maximum driving speed allowed by the road, is a safety reward weight coefficient;
[0043] An efficiency reward function is set to encourage the vehicle to approach the road speed limit, as follows:
[0044]
[0045] is an efficiency reward function, where is an efficiency reward weight coefficient, is the minimum value of the speed allowed by the road;
[0046] A scene-specific reward function is set for the merging lane scene to avoid hindering other vehicles, as follows:
[0047]
[0048] is a merging lane speed penalty reward function, is a weight coefficient thereof, is the target speed of the ego vehicle through the merging lane.
[0049] The present application proposes an automatic driving decision system applied to implement the automatic driving decision method as proposed above, the system comprising:
[0050] an environment perception module for acquiring vehicle state information;
[0051] a trajectory prediction module for initializing a particle set based on the current observed vehicle state, updating the particle set state through a vehicle kinematics model and introducing Gaussian noise to simulate uncertainty, calculating particle weights using a multi-dimensional Gaussian distribution, and outputting a weighted average predicted trajectory after optimization by resampling; the particle weights include position error, heading angle error and speed error;
[0052] a risk assessment module for calculating a longitudinal collision risk value and a lateral collision risk value based on the current position, predicted position, speed and safety distance of the vehicle, and dynamically adjusting the weight according to the traffic density to output a comprehensive driving risk value IDR;
[0053] The decision generation module is used to integrate the comprehensive driving risk value (IDR) and the vehicle driving intention into a dual deep Q network based on attention mechanism enhancement to generate lane-changing decisions for autonomous vehicles.
[0054] The control execution module is used to convert lane-changing decisions into control signals.
[0055] This application proposes a vehicle including an autonomous driving decision-making system as described above.
[0056] The beneficial effects of this invention are as follows:
[0057] This invention achieves high-precision prediction of surrounding vehicle trajectories through a multi-feature Gaussian weighted particle filter algorithm. This algorithm, combined with vehicle kinematic constraints and multi-dimensional error calculation, significantly improves prediction accuracy and environmental adaptability. Based on the prediction results, a comprehensive driving risk assessment model is constructed, establishing a more comprehensive risk perception system by dynamically fusing lateral and longitudinal risk factors, enabling the autonomous driving system to more accurately identify potential hazards. A deep reinforcement learning decision-making algorithm incorporating an attention mechanism effectively focuses on key environmental information, optimizing the information processing efficiency of the decision-making process. The overall technical solution, while ensuring real-time performance, significantly improves the decision-making safety and reliability of autonomous driving systems in complex traffic environments, providing an effective solution to the problems of incomplete risk assessment and low decision-making efficiency in existing autonomous driving technologies. Attached Figure Description
[0058] Figure 1 This is a flowchart of the autonomous driving decision-making method of the present invention;
[0059] Figure 2 This is a flowchart of an autonomous driving decision-making method according to the present invention;
[0060] Figure 3 A detailed flowchart of the particle filter trajectory prediction module;
[0061] Figure 4 This is a schematic diagram of the comprehensive risk assessment module;
[0062] Figure 5 This is a schematic diagram of a simulated environment for a four-lane highway scenario;
[0063] Figure 6 A schematic diagram of a simulated environment for a highway merging lane scenario;
[0064] Figure 7 The collision rate comparison curves are shown for a four-lane highway scenario.
[0065] Figure 8 The average speed comparison curves are shown for a four-lane highway scenario.
[0066] Figure 9 average reward comparison curve for the four-lane highway scenario;
[0067] Figure 10 collision rate comparison curve for the merging lane scenario;
[0068] Figure 11 average speed comparison curve for the merging lane scenario;
[0069] Figure 12 average reward comparison curve for the merging lane scenario;
[0070] Figure 13 driving behavior analysis graph for the four-lane highway scenario;
[0071] Figure 14 driving behavior analysis graph for the merging lane scenario. DETAILED DESCRIPTION
[0072] The following description provides specific applications and requirements of the present specification, in order to enable a person skilled in the art to manufacture and use the content of the present specification. Various partial modifications of the disclosed embodiments are obvious to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of the present specification. Therefore, the present specification is not limited to the embodiments shown, but is consistent with the widest scope of the claims.
[0073] The terms used herein are only for the purpose of describing specific example embodiments, and are not limiting. For example, unless the context clearly indicates otherwise, as used herein, the singular forms "a", "an", and "the" can also include the plural forms. When used in the present specification, the terms "comprise", "include" and / or "contain" mean that the associated integer, step, operation, element and / or component exists, but do not exclude the presence of one or more other features, integers, steps, operations, elements, components and / or groups.
[0074] These features of the present specification and other features, and the operation and function of related elements of the structure, and the economy of combination and manufacture of components can be significantly improved in view of the following description. With reference to the drawings, all of which form part of the present specification. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to limit the scope of the present specification. It should also be understood that the drawings are not drawn to scale.
[0075] The flowcharts used in this specification show operations implemented by systems according to some embodiments in this specification. It should be clearly understood that the operations of the flowcharts can not be implemented in order. Instead, the operations can be implemented in reverse order or simultaneously. In addition, one or more other operations can be added to the flowcharts. One or more operations can be removed from the flowcharts.
[0076] Embodiment 1
[0077] See Figure 1 and Figure 2 , the embodiment proposes an automatic driving decision-making method with dynamic risk perception and attention focusing. By combining the predicted trajectory information, the current state information of the vehicle and the comprehensive risk assessment value, a reasonable lane changing decision scheme is designed to improve the safety and real-time performance of the automatic driving vehicle lane changing decision-making. The method specifically includes the following steps:
[0078] S1. Initialize the particle set based on the current observed vehicle state, update the particle set state by the vehicle kinematics model and introduce Gaussian noise to simulate uncertainty, calculate the particle weight by using multi-dimensional Gaussian distribution, and output the weighted average predicted trajectory after optimization by resampling. The particle weight includes position error, heading angle error and speed error.
[0079] S2. Based on the vehicle's current position, predicted position, speed and safety distance, respectively calculate the longitudinal collision risk value and the lateral collision risk value , and dynamically adjust the weight according to the traffic density to fuse and output the comprehensive driving risk value IDR.
[0080] S3. After integrating the comprehensive driving risk value IDR and the vehicle driving intention, input them into the double deep Q network enhanced based on the attention mechanism to generate the lane changing decision of the automatic driving vehicle.
[0081] It can be understood that the present application realizes intelligent decision-making of automatic driving vehicles through multi-module cooperation. First, trajectory prediction is based on particle filtering algorithm. By initializing the particle set and using vehicle kinematics model to update the state, the particle weight is calculated by combining multi-dimensional Gaussian distribution, and the weighted average predicted trajectory is finally output to provide the prediction ability of future traffic state. Second, risk assessment calculates the longitudinal and lateral collision risk values by real-time analyzing the current state and predicted trajectory of the vehicle, and dynamically adjusts the weight coefficient based on the traffic density to generate the comprehensive driving risk value IDR, realizing the quantitative evaluation of environmental risk. Finally, decision-making integrates the comprehensive risk value and the vehicle state features, and selects features and optimizes decisions through the double deep Q network enhanced by the attention mechanism. The attention mechanism dynamically focuses on key environmental information and suppresses irrelevant interference, while the double deep Q network outputs the optimal lane changing decision based on the reinforcement learning framework under the premise of balancing safety and efficiency.
[0082] Specifically, step S1 includes multi-feature Gaussian weighted particle filter trajectory prediction, and more specifically:
[0083] Particle filter directly approximates posterior probability distribution through Monte Carlo sampling, and thus is suitable for traffic scenes with uncertainty and nonlinear characteristics. Compared with deep learning methods, particle filter does not require a large amount of offline training data, can realize real-time prediction under limited computing resources, and has strong applicability. Considering that the vehicle kinematic model has certain physical constraints, high computational efficiency and high interpretability, the state of the particle is updated according to the vehicle kinematic model. In addition, the weight calculation of the particle is designed based on a multi-dimensional Gaussian distribution, which considers position error, heading angle error and speed error to ensure the accuracy and robustness of weight calculation and improve the reliability of state estimation. The whole prediction process is as shown in Figure 3 .
[0084] (a) Particle initialization
[0085] At the beginning of prediction, a set of particles are randomly generated in the particle state space according to the observed vehicle state (vehicle position, heading angle, speed) in the highway scene. The distribution of these particles should cover the state range of the vehicle as much as possible, and the state space boundary is limited in combination with the prior constraints of the highway scene (such as lane line range, speed limit rules) to ensure the rationality of the particle set.
[0086] (b) Particle update
[0087] The present application realizes particle state prediction based on the vehicle kinematic model, which completely describes the core characteristics of vehicle motion (including position, speed, acceleration and yaw angle, etc. Key state quantities) under the premise of ensuring physical rationality and computational efficiency. The particle update formula designed by the present application is as follows:
[0088]
[0089] wherein are the predicted position, predicted heading angle and predicted speed of the vehicle, respectively; are the current position, current heading angle and current and speed of the vehicle, respectively; is the yaw angle, is the prediction time step, L is the wheelbase of the vehicle, is the Gaussian noise, are the current acceleration of the ego vehicle and the variance of the Gaussian noise, respectively. Based on the above model, the particle is updated, which ensures the interpretability and trajectory smoothness of the vehicle trajectory prediction.
[0090] (c) Weight calculation
[0091] Each particle in particle filter represents a state hypothesis, whose weight is calculated by likelihood function, and the particle weight is proportional to the likelihood function. The likelihood function of the present application is designed based on position error, heading angle error and velocity error. This multi-dimensional weight calculation has stronger anti-interference ability than single error calculation. The likelihood function is as follows:
[0092]
[0093] where respectively represent the error of the observed and predicted values of the vehicle position, heading angle and velocity, respectively are the standard deviations corresponding to the vehicle position, heading angle and velocity. The above formula will calculate and assign a weight value to each particle, which provides the basis for subsequent particle resampling.
[0094] (d) Resampling
[0095] Resampling is performed by selecting particle indices according to the importance weights of the particles, preferentially preserving particles with higher weights. Indices , ,..., are sampled from a categorical distribution parameterized by normalized weights .
[0096] (e) Predicted output
[0097] The predicted vehicle state of the model includes the lateral and longitudinal position of the vehicle, the heading angle and the velocity. The predicted output result is the weighted sum of all particles, and the predicted output value is used as part of the input information of the comprehensive risk assessment model, giving the risk assessment model the future trajectory state information to improve the comprehensiveness of the risk assessment model.
[0098] Specifically, step S2 includes comprehensive risk assessment, more specifically:
[0099] In order to make the autonomous vehicle make real-time accurate decisions, the application designs a comprehensive risk assessment function, which combines the vehicle predicted trajectory to quantify the uncertain risk in the dynamic environment. Unlike traditional rule-based or heuristic methods, the designed method considers a more comprehensive set of vehicle motion states (including the current vehicle state, the vehicle predicted trajectory state and the safety distance factor) to quantitatively assess the lateral and longitudinal risks of the ego vehicle, and dynamically weights and fuses the lateral and longitudinal risks, which can more comprehensively and finely analyze and calculate the potential danger. The model will evaluate the comprehensive driving risk (IDR) of the four vehicles closest to the ego vehicle, and input the quantified risk as the state of the DDQN (Double Deep Q-Network) decision model, thereby enhancing the risk perception ability of the ego vehicle to the traffic situation. The main advantage of this method is its continuous risk estimation, which allows for more smooth and adaptive risk assessment under uncertainty. The comprehensive risk assessment schematic diagram is shown in Figure 4 The longitudinal risk assessment function is defined as follows:
[0100]
[0101]
[0102]
[0103] wherein is the maximum value of the longitudinal risk, is a shape adjustment factor. and represent the longitudinal positions of the surrounding vehicles and the ego vehicle, respectively, and represent the predicted longitudinal positions of the surrounding vehicles and the ego vehicle, respectively. and represent the speed and predicted speed of the ego vehicle and the surrounding vehicles, respectively. and represent the heading angle and predicted heading angle of the ego vehicle and the surrounding vehicles, respectively. , , are the dynamic convergence coefficient and the predicted dynamic convergence coefficient along the longitudinal direction, respectively. Among them is the reference speed, which is set to 25 m / s, is a constant convergence coefficient set to 1.5, is set to 0.8. If the ego vehicle speed and the ego vehicle predicted speed are too high, the dynamic convergence coefficient will be reduced accordingly, which will result in the sensitivity of the risk value to the relative distance change between the ego vehicle and the surrounding vehicles being enhanced, thereby enhancing the response ability of the model to the near distance risk. is the ratio of longitudinal safety distance and real distance. If the real distance between ego vehicle and surrounding vehicle is less than the safety distance, the risk value will be larger, which indicates that the current following or lane changing behavior has a higher collision risk. Conversely, if the real distance is larger than the safety distance, the risk value will be smaller, which indicates that the current driving state is relatively safe. Wherein and are the longitudinal distance and lateral distance between ego vehicle and surrounding vehicle, is the distance from the center of mass of ego vehicle to the front of the vehicle, represents the distance from the center of mass of ego vehicle to the side of the vehicle. is the safety distance margin, which is set to 2 m, is the driver reaction time, which is set to 0.4 s, is the vehicle braking system reaction time, which is set to 0.15 s, is the maximum deceleration of the vehicle, which is set to 10 m / s2; ;
[0104] The lateral risk assessment function has a similar structure to the longitudinal risk assessment function, which is defined as follows:
[0105]
[0106] Wherein is the maximum value of lateral risk, is the lateral safety distance, which is set to 1.5 m based on the road width; the lateral real distance ; and , are the dynamic convergence coefficient and predicted dynamic convergence coefficient in the lateral direction, respectively, wherein is the lateral constant convergence coefficient, which is set to 0.8, is set to 0.5;
[0107] By integrating the longitudinal and lateral risk assessment functions of the ego vehicle and considering the influence of traffic density on lateral and longitudinal risk assessment, the final comprehensive risk assessment function is defined as follows:
[0108]
[0109] In typical road environments, the higher the traffic density, the smaller the vehicle spacing, and the more prone to rear-end collisions, so the longitudinal risk accounts for a larger proportion of the total risk. Therefore, dynamic weights related to traffic density and are designed, and the larger the traffic density, the higher the longitudinal risk. , , is the current traffic density, which is calculated based on the ratio of the number of vehicles within a certain range to the number of roads. and are the minimum and maximum values of the traffic flow density respectively. The final calculated comprehensive driving risk value will be used as the state input of the DDQN decision module, so that the autonomous vehicle has a certain risk awareness ability.
[0110] Specifically, step S3 includes an attention mechanism-based DDQN decision algorithm, and more specifically:
[0111] (a) State processing module based on attention mechanism
[0112] In order to further improve the training efficiency of the agent, the application introduces an attention mechanism into the DDQN decision algorithm to dynamically focus on key vehicle states. This mechanism enables the autonomous vehicle to allocate attention resources according to the current driving situation, more efficiently extract information related to decision-making, and thus improve the effectiveness and generalization ability of policy learning.
[0113] The attention mechanism has certain advantages over traditional feature processing methods such as convolutional neural networks (CNN) or long short-term memory networks (LSTM). Compared with traditional models, the self-attention mechanism can adaptively focus on key vehicle states without being limited by fixed perception ranges, thus improving the relevance of decisions. At the same time, it can model global relationships and capture the potential influence of distant vehicles on current driving decisions, which is very suitable for the needs of autonomous driving to understand complex traffic environments. In addition, it supports parallel computing and does not rely on time series expansion, which is more advantageous in processing efficiency than recurrent neural networks and is suitable for real-time decision-making in autonomous driving.
[0114] The process of using the self-attention mechanism to weight process the vehicle state features in the environment is as follows:
[0115] First, the input state tensor is split into features of the ego vehicle and other vehicles, and the corresponding attention mask is generated.
[0116] Then, the features of the ego vehicle and other vehicles are encoded through a multi-layer perception respectively to obtain embedding representations. The ego vehicle embedding is used as Query, while the concatenation of the ego vehicle and other vehicle embeddings is used as Key and Value. Next, the similarity weight between Query and Key is calculated through the attention mechanism, and the attention output is obtained by combining Value.
[0117] Finally, the output is fused with the original ego vehicle embedding through residual connection to generate updated ego vehicle features for subsequent decision-making tasks based on the DDQN decision algorithm.
[0118] (b) State and action
[0119] In deep reinforcement learning, the state space is the set of all possible states that the agent can perceive when interacting with the environment. The reasonable design of the state space is crucial for the convergence and performance of the algorithm. For the baseline method, the selected state of the environment is: the lateral and longitudinal position of the vehicle, the lateral and longitudinal position deviation, the lateral and longitudinal speed of the vehicle, and the heading angle of the vehicle.
[0120] For the risk-aware DDQN decision framework designed by the present application, a comprehensive driving risk feature is additionally added in the state space. The state feature related to the driving intention of the vehicle. The comprehensive driving risk feature is calculated by formula (7), wherein i={0, 1, 2, 3} represents the serial number of the surrounding vehicle closest to the ego vehicle.
[0121] The present application uses discrete actions as the action space of the DDQN decision algorithm, which includes: left lane change, right lane change, acceleration, keep and deceleration.
[0122] (c) Reward function
[0123] In the reinforcement learning-based autonomous driving decision, the design of the reward function directly affects the learning effect and decision behavior of the agent. Generally, safety and efficiency are the main factors considered in the decision-making process, so the present application considers safety and efficiency factors to design the reward function. However, in the autonomous driving decision-making process, for example, based on the safety distance between the ego vehicle and the preceding vehicle to avoid collision, although the purpose is to ensure driving safety, it may reduce the exploration ability of the agent. Because the agent will focus too much on avoiding collision under such a complex reward rule, and reduce the attempt of other driving strategies, which limits its ability to find the optimal decision strategy in different driving scenarios, therefore a simple and direct reward function is adopted to deal with the vehicle decision problem.
[0124] The safety reward function is defined as follows:
[0125]
[0126] The safety reward function is defined as follows: is the safety reward weight coefficient, is the forward speed of the ego vehicle, is the maximum driving speed allowed on the road, which is set to 30 m / s. If the forward speed of the ego vehicle is faster, the penalty value after the collision will be larger.
[0127] The efficiency reward function is designed as follows:
[0128]
[0129] The efficiency reward function is designed as follows: is an efficiency reward weight coefficient, is a minimum value of road allowed speed, which is set as 20 m / s. The purpose of the speed reward is to encourage the autonomous vehicle to travel within the range of the road allowed speed, and the closer the speed is to the maximum value of the road allowed speed, the higher the reward is. In this way, the agent can learn the strategy of traveling at a higher and appropriate speed.
[0130] For the fusion lane scene built in the application, another reward function is designed, in addition to the above reward function, a merging lane speed penalty reward function is further included, as shown in the following formula:
[0131]
[0132] is a merging lane speed penalty reward function, wherein is a weight coefficient thereof, is a target speed of the ego vehicle through the merging lane, which is set as 30 m / s. The purpose of the reward function is to guide the autonomous vehicle to pass through the merging section at a high speed or to make a reasonable lane change to avoid hindering the merging vehicle.
[0133] Embodiment 2
[0134] The embodiment provides an autonomous driving decision system, which is applied to implement the autonomous driving decision method as proposed in Embodiment 1, and the system comprises:
[0135] an environment perception module, configured to acquire vehicle state information;
[0136] a trajectory prediction module, configured to initialize a particle set based on a currently observed vehicle state, update a particle set state through a vehicle kinematics model and introduce Gaussian noise to simulate uncertainty, calculate particle weights by using a multi-dimensional Gaussian distribution, and output a weighted average predicted trajectory after optimization by resampling; the particle weights comprise a position error, a heading angle error and a speed error;
[0137] a risk assessment module, configured to calculate a longitudinal collision risk value and a lateral collision risk value based on a current position, a predicted position, a speed and a safety distance of the vehicle, and dynamically adjust weights according to a traffic density to fuse and output a comprehensive driving risk value IDR;
[0138] a decision generation module, configured to input the comprehensive driving risk value IDR and a vehicle driving intention after feature integration into a double deep Q network enhanced based on an attention mechanism to generate a lane change decision of the autonomous vehicle;
[0139] a control execution module, configured to convert the lane change decision into a control signal.
[0140] According to the above embodiment, the working principle of the application is to realize intelligent decision of the autonomous vehicle through multi-module cooperation. First, the environment perception module collects the state information of the ego vehicle and surrounding vehicles in real time, including position, heading angle, speed and other key parameters. The trajectory prediction module is based on the particle filter algorithm, which initializes the particle set and updates the state by using the vehicle kinematic model, introduces Gaussian noise to simulate the uncertainty of the traffic environment, calculates the particle weight by using multi-dimensional Gaussian distribution and performs resampling optimization, and finally outputs the weighted average predicted trajectory to provide the prediction ability of the future traffic state for the system. The risk assessment module calculates the longitudinal and lateral collision risk values by analyzing the current state and predicted trajectory of the vehicle in real time, and dynamically adjusts the weight coefficient based on the traffic density to generate the comprehensive driving risk value IDR, realizing the quantitative assessment of the environmental risk. The decision generation module integrates the comprehensive risk value and vehicle state characteristics, and selects features and optimizes decisions through the double deep Q network enhanced by the attention mechanism, wherein the attention mechanism dynamically focuses on key environmental information and suppresses irrelevant interference, and the double deep Q network outputs the optimal lane changing decision under the premise of balancing safety and efficiency based on the reinforcement learning framework. The control execution module converts the decision instruction into vehicle control signals.
[0141] In another embodiment of the application, a vehicle is provided, comprising an autonomous driving decision system as in embodiment 2.
[0142] Specifically, the vehicle described in the application can be a passenger car, a commercial vehicle, a special vehicle or any type of intelligent connected vehicle. The vehicle integrates the autonomous driving decision system of the application in its interior as its "brain" or "decision center", which is closely connected and cooperates with the sensors (such as cameras, lidar, millimeter wave radar) of the vehicle, positioning unit, high-precision map and execution mechanism (such as steering system, driving system, braking system).
[0143] The decision-making method proposed in the application will be further described and explained in combination with specific test cases and some figures.
[0144] In order to study the performance of the proposed decision-making framework, first, the intelligent driver model is used to generate random traffic flow to restore the uncertainty of the real traffic environment. Two common highway traffic scenarios are simulated, in which the road width is 4m, the vehicle length is 5m, and the vehicle width is 1.9m. As shown in Figure 5 and Figure 6The scenarios are shown. Scenario 1 is a four-lane highway scenario with high-density traffic flow, and scenario 2 is a highway merging lane scenario. To reflect the uncertainty of real-world traffic conditions, the positions and speeds of vehicles in all scenarios contain a certain degree of randomness. The observable distance of the autonomous vehicle is 100 meters in front and behind the vehicle. To improve training speed, four parallel training environments are generated, and the decision scheme designed by the present application is compared with four common decision algorithms to illustrate its advantages, including PPO, DQN, DDQN, and A2C, DQN (Deep Q-Network) is a value-based deep reinforcement learning algorithm, A2C (Advantage Actor-Critic) is an Actor-Critic framework algorithm based on policy gradient, PPO (Proximal Policy Optimization) is an improved algorithm based on policy gradient; the training results Figures 7-12
[0145] Figures 7-9 The training results under scenario 1 are shown. Compared with traditional methods, the decision model PDMM (Perception-Decision Model with Multi-feature fusion) proposed by the present application achieves a lower collision rate. However, the average speed of this strategy is also lower than that of other algorithms, and the driving efficiency is slightly reduced, which may be due to the small vehicle spacing in this case, and due to the influence of the risk state, PDMM tends to sacrifice driving speed to improve driving safety. Overall, its average reward is still slightly higher than that of other methods. At the same time, PPO also shows a relatively conservative driving strategy, while DQN, DDQN, and A2C choose a riskier and more efficient driving strategy. The above figure shows that the decision model proposed by the present application can dynamically adjust the decision strategy according to the comprehensive risk assessment value, allowing the agent to perceive the current risk at each decision step, thereby selecting a safer action and effectively reducing the possibility of collision.
[0146] Figures 10-12 The training results in scenario 2 are shown. In this case, PDMM achieves a lower collision rate than other methods while maintaining a relatively high driving speed. It performs well in both safety and driving efficiency, and the average reward is also slightly higher than that of other methods. This indicates that the autonomous vehicle has learned a more efficient and safer decision strategy by integrating risk state assessment and attention mechanism.
[0147] To better understand the impact of environmental uncertainty on the decision-making of autonomous vehicles, the present case analyzes the decision-making scenarios using DDQN and PDMM strategies. Figure 13 is the driving situation under scenario 1, under the DDQN strategy, at the initial moment, since there are front vehicles in the four lanes, the ego vehicle tries to change lanes to overtake at T = 1 s in order to pursue a higher speed reward, but the timing of changing lanes is not appropriate, and at T = 2 s, a collision occurs with D vehicle. Under the PDMM strategy, the ego vehicle has certain risk perception and attention allocation capabilities, and at T = 2 s, the ego vehicle finds the appropriate timing to overtake C vehicle, which shows that the strategy is superior to the DDQN strategy.
[0148] Figure 14 is the driving situation under scenario 2. Under the DDQN strategy, since there are front vehicles A and B on lanes 1 and 2, the ego vehicle wants to continuously change lanes to overtake A and B at T = 1 s, but it does not fully pay attention to the state change of B vehicle during the lane change, and a collision occurs with B vehicle at T = 1.5 s-2 s, which shows that the strategy only considers traffic efficiency and ignores the uncertain risks in the environment. In contrast, PDMM successfully overtakes the front vehicles A and B at T = 1.5 s by predicting vehicle trajectories through particle filtering and combining the comprehensive risk assessment function and the self-attention mechanism to determine the appropriate time for lane changing decision, which ensures both driving safety and driving efficiency.
[0149] In summary, the effectiveness and superiority of the proposed automatic driving decision-making method are fully demonstrated through systematic experimental verification. In simulated complex scenarios such as high-density traffic flow and merging lanes, the method exhibits significant technical advantages: through the innovative particle filtering trajectory prediction algorithm, accurate prediction of the motion trajectories of surrounding vehicles is achieved; the designed comprehensive risk assessment model can comprehensively quantify potential dangers in the traffic environment; and the reinforcement learning decision-making algorithm based on the attention mechanism effectively improves the decision-making efficiency and adaptability of the system. The test results show that the method significantly improves the safety of the autonomous vehicle while ensuring driving efficiency, especially in handling unexpected situations and complex traffic scenarios, and exhibits stronger robustness. Compared with traditional decision-making methods, the present application can more accurately identify risks and more reasonably plan behaviors, enabling autonomous vehicles to make safer and more intelligent decisions in dynamically changing traffic environments.
[0150] The above description is merely the preferred embodiments of the present disclosure and the explanation of the principles of the applied technology. Those skilled in the art should understand that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the disclosed concept. For example, the above features can be replaced with similar functional technical features disclosed in the present disclosure (but not limited to) to form a technical solution.
[0151] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0152] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. An autonomous driving decision-making method with dynamic risk perception and attention focus, characterized in that, Includes the following steps: S1. Initialize the particle set based on the current observed vehicle state, update the particle set state through the vehicle kinematics model and introduce Gaussian noise to simulate uncertainty, calculate the particle weights using a multidimensional Gaussian distribution, and output the weighted average predicted trajectory after resampling optimization; the particle weights include position error, heading angle error and velocity error; S2. Calculate the longitudinal collision risk value based on the vehicle's current position, predicted position, speed, and safe distance. and lateral collision risk value It dynamically adjusts the weights based on traffic density and integrates them to output a comprehensive driving risk value (IDR). S3. The integrated driving risk value (IDR) and vehicle driving intention are integrated and then input into a dual deep Q-network based on attention mechanism enhancement to generate lane-changing decisions for autonomous vehicles.
2. The autonomous driving decision-making method with dynamic risk perception and attention focus according to claim 1, characterized in that, Step S1 includes: S11. Obtain the real-time status data of the observed vehicle at the current moment, including the vehicle's position coordinates, heading angle, and speed; S12. Generate in the three-dimensional state space based on the current observed state data. There are 1 initial particles, and the states of each particle follow a Gaussian distribution with the mean of the observed values; S13. For each particle, calculate the predicted state value at the next moment based on the vehicle kinematics model; S14. For each predicted particle, calculate its three-dimensional error with the latest observation data, including position error, heading angle error, and velocity error; S15. Calculate particle weights based on three-dimensional error, normalize the weights, and perform system resampling to retain high-weight particles; S16. Output the weighted average state of the particles after resampling as the final prediction result.
3. The autonomous driving decision-making method with dynamic risk perception and attention focus according to claim 2, characterized in that, In step S13, the particle state update adopts the following kinematic model: ; in These are the vehicle's predicted position, predicted heading angle, and predicted speed; These are the vehicle's current position, current heading angle, and current speed, respectively. It's the yaw angle. It predicts the time step. L It refers to the vehicle's wheelbase. It's Gaussian noise. These are the variances of the vehicle's current acceleration and Gaussian noise, respectively.
4. The autonomous driving decision-making method with dynamic risk perception and attention focus according to claim 2, characterized in that, Step S2 includes: S21. Obtain the current and predicted states of the vehicle and surrounding vehicles; S22. Calculate longitudinal risk value As shown in the following formula: ; in and These represent the longitudinal positions of the surrounding vehicles and the vehicle itself, respectively. and These represent the predicted longitudinal positions of surrounding vehicles and the vehicle itself, respectively. It is the longitudinal safe distance between the vehicle and surrounding vehicles. It is the actual longitudinal distance between the vehicle and surrounding vehicles; It is the maximum value of vertical risk. It is a shape adjustment factor; , respectively, are the dynamic convergence coefficient along the longitudinal direction and the predicted dynamic convergence coefficient, where This is a reference speed. It is the vehicle's current speed. It is the predicted speed of the vehicle. It is a constant coefficient. It is the longitudinal constant convergence coefficient; S23. Calculate the horizontal risk value As shown in the following formula: ; in It is the maximum value of horizontal risk. This refers to the lateral safe distance between your vehicle and surrounding vehicles, set based on the road width; the actual lateral distance. ,in It is the lateral distance between the vehicle and surrounding vehicles. This represents the distance from the vehicle's center of gravity to the side of the vehicle body; and , , respectively, are the dynamic convergence coefficient along the horizontal direction and the predicted dynamic convergence coefficient, where It is a constant coefficient. It is the horizontal constant convergence coefficient; S24. Calculate the Integrated Risk Value (IDR) as follows: ; in and These are the dynamic weighting coefficients for longitudinal and lateral risks related to traffic density, respectively. , ;in The current traffic density is calculated based on the ratio of the number of vehicles to the number of roads within a certain range; and These are the minimum and maximum values of traffic density, respectively.
5. The autonomous driving decision-making method with dynamic risk perception and attention focus according to claim 4, characterized in that, In step S22, the calculation of the longitudinal safety distance and the actual distance includes the following formula: ; ; in and These represent the speed of the vehicle relative to surrounding vehicles and the predicted speed, respectively. and These represent the heading angle of the vehicle and the predicted heading angle of the surrounding vehicles, respectively. It is the longitudinal distance between the vehicle and surrounding vehicles. It is the distance from the car's center of gravity to the front of the car. It represents the distance from the vehicle's center of gravity to the side of the vehicle body. It is the safety distance margin. It is the driver's reaction time. It is the response time of the vehicle's braking system. It is the vehicle's maximum deceleration.
6. The autonomous driving decision-making method with dynamic risk perception and attention focus according to claim 5, characterized in that, Step S3 includes: S31. Encode the states of the vehicle and surrounding vehicles into feature vectors; S32. Calculate the similarity weights between the vehicle and surrounding vehicles through a self-attention mechanism, and generate environmental features representing the attention distribution by weighted fusion; S33. The features are concatenated with the original vehicle state and input into the DDQN network to generate an action strategy, which includes: left lane change, right lane change, acceleration, deceleration, and holding.
7. The autonomous driving decision-making method with dynamic risk perception and attention focus according to claim 6, characterized in that, The method also includes reward function design, specifically including: The safety reward function is set to be positively correlated with the vehicle's speed in the event of a collision, as shown in the following formula: ; It is a security reward function, where It is the forward speed of the vehicle. It is the maximum speed allowed on the road. This is the safety reward weighting coefficient; an efficiency reward function is set to incentivize vehicles to approach the road speed limit, as shown in the following formula: ; It is an efficiency reward function, where It is the efficiency reward weighting coefficient. It is the minimum speed allowed on the road; Set a scene-specific reward function for merging into the lane to avoid obstructing other vehicles, as shown below: ; It is the merge lane speed penalty and reward function. It is its weighting coefficient. It is the target speed at which the vehicle passes through and merges into the lane.
8. An autonomous driving decision-making system, characterized in that, The system is applied to implement the autonomous driving decision-making method as described in any one of claims 1-7, the system comprising: The environmental perception module is used to acquire vehicle status information; The trajectory prediction module is used to initialize the particle set based on the current observed vehicle state, update the particle set state through the vehicle kinematics model and introduce Gaussian noise to simulate uncertainty, calculate the particle weights using a multidimensional Gaussian distribution, and output the weighted average predicted trajectory after resampling optimization; the particle weights include position error, heading angle error and velocity error; The risk assessment module is used to calculate longitudinal collision risk values based on the vehicle's current position, predicted position, speed, and safe distance. and lateral collision risk value It dynamically adjusts the weights based on traffic density and integrates them to output a comprehensive driving risk value (IDR). The decision generation module is used to integrate the comprehensive driving risk value (IDR) and the vehicle driving intention into a dual deep Q network based on attention mechanism enhancement to generate lane-changing decisions for autonomous vehicles. The control execution module is used to convert lane-changing decisions into control signals.
9. A vehicle, characterized in that, Including an autonomous driving decision-making system as described in claim 8.
Citation Information
Patent Citations
Driving risk prediction method based on multi-level multi-dimensional index system
CN114613127A
Traffic driving safety early warning method and system based on machine learning
CN118762520A