Automatic driving unprotected left turn game decision-making method and system considering interaction style
By establishing a mixture of Gaussian distribution and social level-k game model, combined with the Monte Carlo tree search algorithm, the unprotected left-turn decision of autonomous vehicles is optimized, solving the complexity problem of vehicle interaction in left turns at unprotected intersections and achieving more efficient and safer traffic.
Patent Information
- Application Number
- CN202511349030.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Existing autonomous driving decision-making methods for left turns at intersections without protection ignore the complexity of traffic, resulting in overly conservative approaches and weak generalization, making it difficult to achieve efficient and safe vehicle interaction.
We adopt an autonomous driving unprotected left-turn game-theoretic decision-making method that considers interaction style. By acquiring vehicle interaction features, we establish a mixture Gaussian distribution model, use Bayesian inference to update the interaction style confidence, and combine a social level-k game model and Monte Carlo tree search algorithm to optimize vehicle decision-making.
It achieves a higher success rate, longer average travel distance, higher average travel speed, and smaller rate of acceleration change in complex intersection scenarios, thereby improving the interactive safety and traffic efficiency of autonomous vehicles.
Smart Images

Figure CN120840666A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of vehicle driving decision-making technology, and specifically relates to an unprotected left-turn game-theoretic decision-making method and system for autonomous driving that takes into account interaction style. Background Technology
[0002] Among various driving environments, urban intersections are one of the hot topics and pain points in autonomous driving research. As a typical complex, dynamic, and time-varying scenario, research on decision-making and planning for crossing behavior at urban intersections is helpful for the implementation of autonomous driving. Left turns at unprotected intersections, as a special type of intersection traffic behavior, are among the most complex autonomous driving scenarios due to the shared traffic area for both straight-going and left-turning vehicles and the lack of dedicated left-turn traffic lights. The interaction between left-turning vehicles and oncoming straight-going vehicles during the crossing process poses challenges to decision-making and planning.
[0003] In existing technologies, rule-based decision-making methods are typically used for left turns at unprotected intersections. While rule-based decision-making methods are logically simple, easy to implement, and can directly draw on human driving experience, they often ignore the complexity of traffic and are usually too conservative and have weak generalization ability. Summary of the Invention
[0004] To address the aforementioned technical issues, this invention proposes an unprotected left-turn game-theoretic decision-making method and system for autonomous driving that considers interaction style, ensuring both interaction safety and improved traffic efficiency.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: A game-theoretic decision-making method for unprotected left turns in autonomous driving, considering interaction styles, includes the following steps: The temporal and dynamic interaction features between the target vehicle and other vehicles when the target vehicle turns left are obtained, and both temporal and dynamic interaction features are normalized. The time interaction parameters are calculated using normalized time interaction features, and the dynamic interaction parameters are calculated using normalized dynamic interaction features; then the comprehensive interaction parameters are calculated using the time interaction parameters and the dynamic interaction parameters. The distribution of interaction styles is modeled as a mixture Gaussian distribution, and based on Bayesian inference, the confidence of other vehicles' interaction styles is updated in real time using the comprehensive interaction parameters. A social level-k game model is established, and the confidence of the updated interaction style of other vehicles is used as the weight to calculate the weighted sum of the expected rewards under different interaction styles, so as to obtain the comprehensive reward function. Based on a predefined discrete action space, the vehicle's state trajectory in the future time domain is predicted using a vehicle kinematics model. A Monte Carlo tree search algorithm incorporating comfort constraints is employed to solve the problem with the comprehensive reward function as the optimization objective, outputting the optimal decision action sequence for the target vehicle.
[0006] This invention also proposes an unprotected left-turn game decision-making method for autonomous driving that considers interaction style, including a data acquisition module, a calculation module, an update module, a model building module, and a solution module; The data acquisition module is used to acquire the temporal and dynamic interaction features between the target vehicle and other vehicles when the target vehicle turns left, and to normalize both the temporal and dynamic interaction features. The calculation module is used to calculate the time interaction parameters using the normalized time interaction features and the dynamic interaction parameters using the normalized dynamic interaction features; then, it uses the time interaction parameters and the dynamic interaction parameters to calculate the comprehensive interaction parameters. The update module is used to model the distribution of interaction styles as a mixture Gaussian distribution, and based on Bayesian inference, update the confidence of other vehicles' interaction styles in real time through the comprehensive interaction parameters. The model building module is used to build a social level-k game model, using the confidence of the updated interaction style of other vehicles as weights, and calculating the weighted sum of expected rewards under different interaction styles to obtain a comprehensive reward function. The solution module is used to predict the state trajectory of the vehicle in the future time domain based on a predefined discrete action space and a vehicle kinematics model. It employs a Monte Carlo tree search algorithm that incorporates comfort constraints, and uses the comprehensive reward function as the optimization objective to solve the problem, outputting the optimal decision action sequence of the target vehicle.
[0007] The effects described in the invention are merely those of the embodiments, and not all the effects of the invention. One of the above technical solutions has the following advantages or beneficial effects: This invention proposes an unprotected left-turn game decision-making method and system for autonomous driving that considers interaction style, belonging to the field of vehicle driving decision-making technology. The method includes the following steps: obtaining the temporal interaction features and dynamic interaction features between the target vehicle and other vehicles when the target vehicle turns left, and normalizing both the temporal interaction features and the dynamic interaction features. Temporal interaction parameters are calculated using normalized temporal interaction features, and dynamic interaction parameters are calculated using normalized dynamic interaction features. A comprehensive interaction parameter is then calculated using the temporal and dynamic interaction parameters. The distribution of interaction styles is modeled as a Gaussian mixture distribution, and based on Bayesian inference, the confidence level of other vehicles' interaction styles is updated in real time using the comprehensive interaction parameters. A social level-k game model is established, using the updated confidence level of other vehicles' interaction styles as weights to calculate the weighted sum of expected rewards under different interaction styles, thus obtaining the comprehensive reward function. Based on a predefined discrete action space, the vehicle's state trajectory in the future time domain is predicted using a vehicle kinematics model. A Monte Carlo tree search algorithm incorporating comfort constraints is used to solve the problem with the comprehensive reward function as the optimization objective, outputting the optimal decision action sequence for the target vehicle. Based on the game-theoretic decision-making method for unprotected left turns in autonomous driving that considers interaction styles, a game-theoretic decision-making system for unprotected left turns in autonomous driving that considers interaction styles is also proposed. This invention addresses the task of turning left at unsignalized intersections by using a level-k game model to model the interactive game relationship between autonomous vehicles and human-driven vehicles. Furthermore, it utilizes a proposed interaction style recognition module to identify the distribution of interaction styles of surrounding vehicles, thereby enhancing the social and interactive nature of the game model.
[0008] This invention can accurately estimate the interaction styles of surrounding traffic participants through real-time behavior observation, enabling autonomous vehicles to exhibit human-like decision-making logic in complex intersection scenarios. This ensures both interaction safety and improves traffic efficiency. Compared with existing technologies, it can achieve a higher success rate, a longer average travel distance, a higher average travel speed, and a smaller average acceleration change rate. Attached Figure Description
[0009] Figure 1 This is a schematic diagram illustrating the implementation of the unprotected left-turn game-theoretic decision-making method for autonomous driving that considers interactive styles in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of a three-car game scenario at an intersection, as proposed in Embodiment 1 of the present invention. Figure 3 This is a schematic diagram of the unprotected left-turn game decision-making system for autonomous driving that considers interactive style, as proposed in Embodiment 2 of the present invention. Detailed Implementation
[0010] To clearly illustrate the technical features of this solution, the invention will be described in detail below through specific embodiments and in conjunction with the accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, components and arrangements of specific examples are described below. Furthermore, reference numerals and / or letters may be repeated in different examples. This repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. Descriptions of well-known components and processing techniques and processes are omitted to avoid unnecessarily limiting the invention.
[0011] Example 1 Embodiment 1 of this invention proposes an unprotected left-turn game decision-making method for autonomous driving that considers interaction style, in order to solve the technical problems of existing technologies that ignore the complexity of traffic, are too conservative, and have weak generalization. Figure 1 This is a schematic diagram illustrating the implementation of the unprotected left-turn game-theoretic decision-making method for autonomous driving that considers interactive styles in Embodiment 1 of the present invention; The system acquires real-time status information of the vehicle and the oncoming vehicles, including position, speed, acceleration, and heading angle. Select the temporal and dynamic interaction features between the target vehicle and other vehicles when the target vehicle turns left, and normalize both the temporal and dynamic interaction features. Temporal interaction features include collision time and collision time change rate; dynamic interaction features include other vehicle speed and other vehicle acceleration.
[0012] To ensure consistency in the scale of interaction features, both temporal and dynamic interaction features are normalized, specifically as follows: ; ; ; ; in, This refers to the normalized collision time. This represents the normalized rate of change of collision time. The normalized speed of the other vehicle; The normalized acceleration of the other vehicle; The parameters that determine the steepness of the collision time curve; The parameter that determines the steepness of the rate of change of the collision time; Parameters used to determine the steepness of the slope for other vehicles; The parameters that determine the steepness of the vehicle's acceleration are: the larger the value of the parameter that determines the steepness, the more sensitive the response. This represents the offset of the collision time. This is the offset of the rate of change of the collision time; The offset of the other vehicle's speed; The offset of the acceleration of the other vehicle.
[0013] The parameter values and offsets that determine the steepness can be determined by analyzing the distribution of data from a real dataset. This can be done by observing the extracted velocity and acceleration data. and The smaller the value, the stronger the tendency to yield. and The larger the value, the stronger the tendency to yield.
[0014] To fully consider the combined effect of each interaction feature, a temporal interaction parameter (TIP) defined by the temporal interaction feature and a dynamic interaction parameter (DIP) defined by the dynamic interaction feature were designed, as follows: ; ; in, For time-interaction parameters; For dynamic interactive parameters; ; in, For comprehensive interactive parameters.
[0015] Interaction parameters right and Take the geometric mean to balance the effects of both and ensure the result is within [0,1]. The closer the value is to 1, the stronger the tendency to yield; the closer it is to 0, the stronger the tendency to cut in front.
[0016] In complex scenarios like unsignalized intersections, the social interaction styles of other traffic participants are constantly changing. Ignoring the uncertainty of these interaction styles can lead to inappropriate or overly conservative decisions. Therefore, this invention proposes a social level-k game model that utilizes comprehensive interaction parameters. The interaction styles of interactive objects are estimated, including conservative, natural, and aggressive styles. Bayesian inference is used to update the confidence scores for each interaction style in real time. The confidence scores for aggressive, conservative, and natural styles are defined as follows: Level-0 is defined as an aggressive interaction style, level-1 is defined as a conservative interaction style, and level-2 is defined as an interaction style with an aggressiveness level between level 0 and level 1, i.e., a natural interaction style.
[0017] Then, the confidence scores of different interaction styles are used to construct a level-k game model. Specifically, the distribution of interaction styles is modeled as a Gaussian mixture distribution, and based on Bayesian inference, the confidence scores of other vehicles' interaction styles are updated in real time using the comprehensive interaction parameters.
[0018] The distribution of interaction styles is modeled as a Gaussian mixture distribution. The Gaussian mixture distribution is used to address the inherent uncertainty and complexity associated with the multimodal interaction styles of interactive objects. It expresses uncertainty with multiple modes by integrating multiple continuous probability distributions.
[0019] The mixture Gaussian distribution is represented as: ; in, It is a mixture Gaussian distribution function; It is the weight of each interaction style and , It is the probability density function of the distribution of each interaction style, with a mean of covariance is .
[0020] The interaction style of interactive objects is constantly changing and needs to be updated in real time. Bayesian inference is used to determine the confidence level for these real-time updates. ; in, j are both interaction style indices; It is Posterior probability of each interaction style; It is The prior probability of each interaction style; It is The prior probability of each interaction style; It is The probability density function of each interaction style distribution, and in An assessment will be conducted at the site.
[0021] The social level-k game model is established as follows: ; ; ; It is Confidence level of each interaction style; For the first One interaction style; The social level-k model can comprehensively consider the distribution of different social interaction styles of interactive objects and can dynamically update the confidence of interaction styles.
[0022] A social level-k game model is established, and the confidence of the updated interaction style is used as the weight to calculate the weighted sum of the expected rewards under different interaction styles, so as to obtain the comprehensive reward function.
[0023] Based on a predefined discrete action space, the vehicle's state trajectory in the future time domain is predicted using a vehicle kinematics model.
[0024] The vehicle kinematics model is used to determine the current time. state vector and the selected action Calculate the next time step state vector Specifically: ; in, for Vehicles at all times coordinate; for Vehicles at all times coordinate; for The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; for Vehicles at all times coordinate; for Vehicles at all times coordinate; for The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; For time step; The predefined discrete action space includes: Keep:( =(0,0); Turn left: ( )=( );Turn right:( )=( ); accelerate:( =(2.5,0); Deceleration: ( =(-2.5,0); Emergency braking: ( )=(-5,0); Quick left turn: ( )=( Quick right turn: ( )=( ).
[0025] The reward for choosing a strategy is defined through rolling time-domain optimization. Indicates the vehicle in the time domain The vehicle selects a sequence of actions. The vehicle's goal is to find the sequence that yields the best overall reward. The largest optimal action sequence and execute the first action in the action sequence. This process will be repeated at every subsequent moment. The comprehensive reward function is defined as follows: The comprehensive reward function is: ; ; ; in, This represents the discount factor for the vehicle. and The meanings are the same; This indicates that the vehicle takes the optimal sequence of actions from other vehicles. In this situation, the vehicle takes action. The instant reward received; Represents the time domain; This represents the optimal sequence of actions for the vehicle; it indicates that the optimal actions of the vehicle are obtained by relying on the optimal actions of other vehicles, which reflects the interaction between agents.
[0026] To better guide vehicles in selecting optimal strategies, instant rewards are provided. Taking into account multiple factors, it is defined as: ; Rewards are divided into three types: security rewards Efficiency Passage Reward and comfort bonus . , as well as These are the weighting coefficients for the three reward types.
[0027] Safety Rewards It can be represented as follows: ; in This is to prevent vehicles from violating traffic regulations, such as crossing solid lane lines or leaving the drivable area. When performing this action... Subsequently, if a traffic violation occurs, a substantial penalty will be imposed. . The definition is as follows: ; Penalties for maintaining safe distances, in order to guide vehicles to maintain a safe distance: ; For vehicle spacing, For safe distance, This is a constant coefficient. If the distance between vehicles is less than the safe distance, a penalty is returned based on the difference between the vehicle spacing and the safe distance.
[0028] Efficiency Pass Rewards To guide intelligent agents through intersection areas quickly and improve traffic efficiency. It can be represented as follows: ; in This indicates a reward for keeping the vehicle as close to the reference speed as possible. Defined as: ; This is a reference speed. It is a constant coefficient. Designed to be the speed of the vehicle A greater penalty is imposed when the value is less than 0, therefore It should be set to a larger coefficient.
[0029] The reward is related to the distance between the vehicle and the target location. Defined as: ; in It refers to the vehicle's position and heading angle. This represents the position and heading angle of the target point.
[0030] Comfort Reward To avoid drastic acceleration or deceleration changes during vehicle operation, the rate of change of acceleration is selected. To build : ; It is the rate of change of acceleration. It is a constant coefficient.
[0031] After constructing the level-k game model, the Monte Carlo tree search method was used to solve the game. The Monte Carlo tree search algorithm consists of the following parts: selection, expansion, simulation, and backtracking.
[0032] Comfort constraints are used to limit the action expansion and simulation process in the Monte Carlo tree search algorithm; the selection phase usually uses the Upper Confidence Bound (UCB) formula to balance exploration and utilization, that is, to consider both good nodes that have been discovered and nodes that have not been fully explored; in the expansion phase, actions that meet the comfort conditions are expanded into child nodes. ; in, It is The average reward of each node It is a constant used to balance exploration and exploitation. It represents the total number of explorations. This is the number of times this node has been accessed. This is the cumulative reward for this node. By calculating the UCB value of each child node and selecting the child node with the highest value, the search process can be effectively guided.
[0033] During the simulation phase, a strategy that satisfies comfort conditions is used for action selection; the comfort conditions are whether the rate of change of acceleration exceeds a threshold.
[0034] A stochastic strategy is typically used for action selection. This strategy is relatively simple and provides sufficient information to estimate the potential reward of a node. However, the stochastic strategy does not consider the rationality of the selected action; the action selected by the stochastic strategy may not meet the actual requirements, which is unreasonable. Therefore, the stochastic strategy must consider comfort requirements. Similarly, in the expansion phase, only actions that meet the above comfort requirements are expanded into child nodes at leaf nodes, not all actions. It is a function that determines whether the action meets comfort requirements. If the action does not meet comfort requirements, the function is true; otherwise, it is false. ; The acceleration of the action at time t+1, The acceleration of the action at time t, This is the limit value for jerk.
[0035] The autonomous driving unprotected left-turn game decision-making method proposed in Embodiment 1 of this invention considers interaction styles. For the left-turn task at an unsignalized intersection, it uses a level-k game model to model the interaction game relationship between autonomous vehicles and human-driven vehicles, and uses the proposed interaction style recognition module to identify the distribution of interaction styles of surrounding vehicles, thereby enhancing the social and interactive nature of the game model.
[0036] To fully verify the execution process of the unprotected left-turn game decision-making method for autonomous driving that considers interaction style proposed in Embodiment 1 of the present invention, Figure 2 This is a schematic diagram of a three-car game scenario at an intersection, as proposed in Embodiment 1 of the present invention. The blue car represents the driver, who is about to turn left. The red and green cars represent the other cars, which are going straight. Each car is controlled by the method proposed in this invention. Taking the driver as an example, the driver first calculates the collision time, the rate of change of the collision time, the speed of the other cars, and their acceleration based on sensor data. Then, it calculates the interaction parameters of the two other cars based on the above data. Subsequently, based on their respective interaction parameters, the distribution of the interaction style of each other car is updated using Bayesian inference. Finally, Monte Carlo tree search is used to find the optimal strategy for the driver. In this example, the driver can safely pass between the two other cars and complete the left turn.
[0037] The autonomous driving unprotected left-turn game decision-making method proposed in Embodiment 1 of this invention, which considers interaction style, can accurately estimate the interaction style of surrounding traffic participants through real-time behavior observation, enabling autonomous vehicles to exhibit human-like decision-making logic in complex intersection scenarios. This ensures both interaction safety and improves traffic efficiency. Compared with existing technologies, it can achieve a higher success rate, a longer average travel distance, a higher average travel speed, and a smaller average acceleration change rate.
[0038] Example 2 Based on the autonomous driving unprotected left-turn game decision-making method considering interaction style proposed in Embodiment 1 of this invention, Embodiment 2 of this invention also proposes an autonomous driving unprotected left-turn game decision-making system considering interaction style. Figure 3 This is a schematic diagram of the unprotected left-turn game decision-making system for autonomous driving that considers the interaction style proposed in Embodiment 2 of the present invention. The system includes: a data acquisition module, a calculation module, an update module, a model building module, and a solution module. The data acquisition module is used to acquire the temporal and dynamic interaction features between the target vehicle and other vehicles when the target vehicle turns left, and to normalize both the temporal and dynamic interaction features. The calculation module is used to calculate the time interaction parameters using normalized time interaction features and the dynamic interaction parameters using normalized dynamic interaction features; then, the comprehensive interaction parameters are calculated using the time interaction parameters and the dynamic interaction parameters. The update module is used to model the distribution of interaction styles as a mixture Gaussian distribution, and based on Bayesian inference, update the confidence of other vehicles' interaction styles in real time through the comprehensive interaction parameters. The model building module is used to build a social level-k game model. It uses the confidence of the updated interaction style of other vehicles as weights to calculate the weighted sum of expected rewards under different interaction styles in order to obtain the comprehensive reward function. The solution module is used to predict the state trajectory of the vehicle in the future time domain based on a predefined discrete action space and a vehicle kinematic model. It employs a Monte Carlo tree search algorithm that incorporates comfort constraints, and uses the comprehensive reward function as the optimization objective to solve the problem, outputting the optimal decision action sequence of the target vehicle.
[0039] In the data acquisition module of this application, the time interaction features include collision time and collision time change rate; the dynamic interaction features include the speed of other vehicles and the acceleration of other vehicles.
[0040] In the calculation module, both temporal interaction features and dynamic interaction features are normalized, specifically as follows: ; ; ; ; in, This refers to the normalized collision time. This represents the normalized rate of change of collision time. The normalized speed of the other vehicle; The normalized acceleration of the other vehicle; The parameters that determine the steepness of the collision time curve; The parameter that determines the steepness of the rate of change of the collision time; Parameters used to determine the steepness of the slope for other vehicles; The parameters used to determine the steepness of the acceleration of other vehicles; This represents the offset of the collision time. This is the offset of the rate of change of the collision time; The offset of the other vehicle's speed; The offset of the acceleration of the other vehicle.
[0041] The time interaction parameters are calculated using normalized time interaction features, and the dynamic interaction parameters are calculated using normalized dynamic interaction features. Finally, the comprehensive interaction parameters are calculated using both the time interaction parameters and the dynamic interaction parameters. Specifically: ; ; in, For time-interaction parameters; For dynamic interactive parameters; ; in, For comprehensive interactive parameters.
[0042] In the update module, the Gaussian mixture distribution is represented as: ; in, It is a mixture Gaussian distribution function; It is the weight of each interaction style and , It is the probability density function of the distribution of each interaction style, with a mean of covariance is ; Confidence level for real-time updates of other vehicles' interaction styles using Bayesian inference: ; in, j are both interaction style indices; It is Posterior probability of each interaction style; It is The prior probability of each interaction style; It is The prior probability of each interaction style; It is The probability density function of each interaction style distribution, and in An assessment will be conducted at the site.
[0043] The interaction styles include aggressive, conservative, and natural; and the interaction styles correspond to the levels of the social level-k game, with aggressive corresponding to level-0, conservative to level-1, and natural to level-2; the Bayesian inference is used to update the confidence of each interaction style based on the comprehensive interaction parameters observed at the current moment.
[0044] In the model building module, the social level-k game model is built as follows: ; ; It is Confidence level of each interaction style; The comprehensive reward function is: ; ; ; in, This represents the discount factor for the vehicle. and The meanings are the same; This indicates that the vehicle takes the optimal sequence of actions from other vehicles. In this situation, the vehicle takes action. The instant reward received; Represents the time domain; This represents the optimal sequence of actions for the vehicle; it indicates that the optimal actions of the vehicle are obtained by relying on the optimal actions of other vehicles, which reflects the interaction between agents.
[0045] In the solution module, the vehicle kinematics model is used based on the current time. state vector and the selected action Calculate the next time step state vector Specifically: ; in, for Vehicles at all times coordinate; for Vehicles at all times coordinate; for The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; for Vehicles at all times coordinate; for Vehicles at all times coordinate; for The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; For time step.
[0046] The predefined discrete action space includes: (Maintaining: ( =(0,0); Turn left: ( )=( );Turn right:( )=( );accelerate:( =(2.5,0); Deceleration: ( =(-2.5,0); Emergency braking: ( = (-5, 0); Quick left turn: ( )=( Quick right turn: ( )=( ).
[0047] A Monte Carlo tree search algorithm incorporating comfort constraints is employed, with the comprehensive reward function as the optimization objective, to solve for the optimal decision-making sequence of the target vehicle, specifically: Comfort constraints are used to limit the action expansion and simulation process in the Monte Carlo tree search algorithm; During the expansion phase, actions that satisfy comfort conditions are expanded into child nodes; During the simulation phase, a strategy that satisfies comfort conditions is used for action selection; the comfort conditions are whether the rate of change of acceleration exceeds a threshold.
[0048] The autonomous driving unprotected left-turn game decision system proposed in Embodiment 2 of this invention considers interaction styles. For the left-turn task at an unsignalized intersection, it uses a level-k game model to model the interaction game relationship between autonomous vehicles and human-driven vehicles, and uses the proposed interaction style recognition module to identify the distribution of interaction styles of surrounding vehicles, thereby enhancing the social and interactive nature of the game model.
[0049] The autonomous driving unprotected left-turn game decision-making system proposed in Embodiment 2 of this invention, which considers interaction style, can accurately estimate the interaction style of surrounding traffic participants through real-time behavior observation, enabling autonomous vehicles to exhibit human-like decision-making logic in complex intersection scenarios. This ensures both interaction safety and improves traffic efficiency. Compared with existing technologies, it can achieve a higher success rate, a longer average travel distance, a higher average travel speed, and a smaller average acceleration change rate.
[0050] The description of the relevant parts of the autonomous driving unprotected left turn game decision system considering interaction style provided in Embodiment 2 of this application can be found in the detailed description of the corresponding parts of the autonomous driving unprotected left turn game decision method considering interaction style provided in Embodiment 1 of this application, and will not be repeated here.
[0051] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that the elements inherent in a process, method, article, or apparatus that includes a list of elements are included. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. Additionally, portions of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of corresponding technical solutions in the prior art have not been described in detail to avoid excessive elaboration.
[0052] While specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art can make other modifications or variations based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A game-theoretic decision-making method for unprotected left turns in autonomous driving, considering interaction styles, characterized in that: The following steps are involved: The temporal and dynamic interaction features between the target vehicle and other vehicles when the target vehicle turns left are obtained, and both temporal and dynamic interaction features are normalized. The time interaction parameters are calculated using normalized time interaction features, and the dynamic interaction parameters are calculated using normalized dynamic interaction features; then the comprehensive interaction parameters are calculated using the time interaction parameters and the dynamic interaction parameters. The distribution of interaction styles is modeled as a mixture Gaussian distribution, and based on Bayesian inference, the confidence of other vehicles' interaction styles is updated in real time using the comprehensive interaction parameters. A social level-k game model is established, and the confidence of the updated interaction style of other vehicles is used as the weight to calculate the weighted sum of the expected rewards under different interaction styles, so as to obtain the comprehensive reward function. Based on a predefined discrete action space, the vehicle's state trajectory in the future time domain is predicted using a vehicle kinematics model. A Monte Carlo tree search algorithm incorporating comfort constraints is employed to solve the problem with the comprehensive reward function as the optimization objective, outputting the optimal decision action sequence for the target vehicle.
2. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 1, characterized in that, The time-interaction features include collision time and collision time change rate; the dynamic interaction features include the speed of other vehicles and the acceleration of other vehicles.
3. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 2, characterized in that, Both temporal and dynamic interaction features were normalized, specifically as follows: ; ; ; ; in, This refers to the normalized collision time. This represents the normalized rate of change of collision time. The normalized speed of the other vehicle; The normalized acceleration of the other vehicle; The parameters that determine the steepness of the collision time curve; The parameter that determines the steepness of the rate of change of the collision time; Parameters used to determine the steepness of the slope for other vehicles; Parameters used to determine the steepness of acceleration of other vehicles; This represents the offset of the collision time. This is the offset of the rate of change of the collision time; The offset of the other vehicle's speed; The offset of the acceleration of the other vehicle.
4. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 3, characterized in that, The time interaction parameters are calculated using normalized time interaction features, and the dynamic interaction parameters are calculated using normalized dynamic interaction features. Finally, the comprehensive interaction parameters are calculated using both the time interaction parameters and the dynamic interaction parameters. Specifically: ; ; in, For time-interaction parameters; For dynamic interactive parameters; ; in, For comprehensive interactive parameters.
5. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 4, characterized in that, The distribution of interaction styles is modeled as a Gaussian mixture distribution. Based on Bayesian inference, the confidence level of other vehicles' interaction styles is updated in real time using the comprehensive interaction parameters. Specifically: The mixture Gaussian distribution is represented as: ; in, It is a mixture Gaussian distribution function; It is the weight of each interaction style and , It is the probability density function of the distribution of each interaction style, with a mean of The covariance is ; Confidence level for real-time updates of other vehicles' interaction styles using Bayesian inference: ; in, j are both interaction style indices; It is the first Posterior probability of each interaction style; It is the first The prior probability of each interaction style; It is the first The prior probability of each interaction style; It is the first The probability density function of each interaction style distribution, and in An assessment will be conducted at the site.
6. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 5, characterized in that, The interaction styles include aggressive, conservative, and natural; and the interaction styles correspond to the levels of the social level-k game, with aggressive corresponding to level-0, conservative to level-1, and natural to level-2; the Bayesian inference is used to update the confidence of each interaction style based on the comprehensive interaction parameters observed at the current moment.
7. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 6, characterized in that, The social level-k game model is established as follows: ; ; ; It is the first Confidence level of each interaction style; For the first One interaction style; The comprehensive reward function is: ; ; ; in, This represents the discount factor for the vehicle. and The meanings are the same; This indicates that the vehicle takes the optimal sequence of actions from other vehicles. In this situation, the vehicle takes action. The instant reward received; This represents the optimal sequence of actions for the vehicle. Represents the time domain.
8. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 1, characterized in that, The vehicle kinematics model is used to determine the current time. state vector and the selected action Calculate the next time step state vector Specifically: ; in, for Vehicles at all times coordinate; for Vehicles at all times coordinate; For The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; for Vehicles at all times coordinate; for Vehicles at all times coordinate; For The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; For time step; The predefined discrete action space includes: (Maintaining: ( =(0,0); Turn left: ( )=( );Turn right:( )=( );accelerate:( =(2.5,0); Deceleration: ( =(-2.5,0); Emergency braking: ( = (-5, 0); Quick left turn: ( )=( Quick right turn: ( )=( ).
9. The unprotected left-turn game-theoretic decision-making method for autonomous driving considering interaction style as described in claim 1, characterized in that, A Monte Carlo tree search algorithm incorporating comfort constraints is employed to solve for the optimal decision-making sequence of the target vehicle, using the comprehensive reward function as the optimization objective. Specifically, the optimal sequence of decision-making actions for the target vehicle is output as follows: Comfort constraints are used to limit the action expansion and simulation process in the Monte Carlo tree search algorithm; During the expansion phase, actions that satisfy comfort conditions are expanded into child nodes; During the simulation phase, a strategy that satisfies comfort conditions is used for action selection; the comfort conditions are whether the rate of change of acceleration exceeds a threshold.
10. A game-theoretic decision-making method for unprotected left turns in autonomous driving, considering interaction styles, characterized in that: It includes a data acquisition module, a calculation module, an update module, a model building module, and a solution module; The data acquisition module is used to acquire the temporal and dynamic interaction features between the target vehicle and other vehicles when the target vehicle turns left, and to normalize both the temporal and dynamic interaction features. The calculation module is used to calculate the time interaction parameters using the normalized time interaction features and the dynamic interaction parameters using the normalized dynamic interaction features; then, it uses the time interaction parameters and the dynamic interaction parameters to calculate the comprehensive interaction parameters. The update module is used to model the distribution of interaction styles as a mixture Gaussian distribution, and based on Bayesian inference, update the confidence of other vehicles' interaction styles in real time through the comprehensive interaction parameters. The model building module is used to build a social level-k game model, using the confidence of the updated interaction style of other vehicles as weights, and calculating the weighted sum of expected rewards under different interaction styles to obtain a comprehensive reward function. The solution module is used to predict the state trajectory of the vehicle in the future time domain based on a predefined discrete action space and a vehicle kinematics model. It employs a Monte Carlo tree search algorithm that incorporates comfort constraints, and uses the comprehensive reward function as the optimization objective to solve the problem, outputting the optimal decision action sequence of the target vehicle.
Citation Information
Patent Citations
An intelligent vehicle driving decision method based on generative countermeasure network
CN109131348A
Intelligent vehicle behavior decision-making method, planning method, system and storage medium
CN114919578A
Intelligent driving decision-making method, decision-making device and vehicle
CN115503756A
Unprotected intersection passing method and device, computer equipment and storage medium
CN119116949A
Game theoretic decision making
US20230182014A1