Automatic driving unprotected left turn game decision method and system considering interaction style

By establishing a mixture of Gaussian distribution and social level-k game model, the unprotected left turn decision of autonomous vehicles is optimized, solving the problems of conservative decision-making and weak generalization in existing technologies, and realizing efficient and safe interactive behavior.

CN120840666BActive Publication Date: 2025-12-09ADVANCED TECH RES INST OF BEIJING UNIV OF TECH +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511349030.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-09
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing autonomous driving decision-making methods for left turns at intersections without protection ignore the complexity of traffic, resulting in overly conservative decisions with weak generalization, making it difficult to achieve efficient and safe interactive behavior.

Method used

We adopt an autonomous driving unprotected left-turn game-theoretic decision-making method that considers interaction style. By acquiring vehicle interaction features, we establish a mixture Gaussian distribution model, use Bayesian inference to update the interaction style confidence, and combine a social level-k game model and Monte Carlo tree search algorithm to optimize vehicle decision actions.

Benefits of technology

It enables efficient and safe vehicle interaction in complex intersection scenarios, improving the success rate of passage, average passage distance and speed, while reducing the rate of change of acceleration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120840666B_ABST
    Figure CN120840666B_ABST
Patent Text Reader

Abstract

The application provides an automatic driving unprotected left turn game decision method and system considering interactive style, and belongs to the technical field of vehicle driving decision.The method is as follows: time interaction features and dynamic interaction features of a target vehicle during left turn are acquired; time interaction parameters and dynamic interaction parameters are respectively calculated by using normalized data; comprehensive interaction parameters are determined; the confidence of the interactive style of another vehicle is updated by Bayesian inference based on the distribution of the comprehensive interaction parameters; a game model is established, the weighted sum of expected rewards under different interactive styles is calculated by taking the confidence of the updated interactive style of another vehicle as the weight to obtain a comprehensive reward function; the state trajectory of the vehicle in a future time domain is predicted by using a vehicle kinematics model based on a pre-defined action space, and the optimal decision action sequence is output by solving the comprehensive reward function as an optimization target.The corresponding system is also provided based on the method.The application guarantees the interaction safety and improves the traffic efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of vehicle driving decision-making, and particularly relates to an automatic driving unprotected left turn game decision-making method and system considering interaction style. BACKGROUND

[0002] Among many driving environments, urban intersection is one of the hot and painful points of automatic driving research. As a typical complex dynamic time-varying scene, the decision and planning research on its crossing behavior is helpful to the landing of automatic driving behavior. Unprotected intersection left turn as a special intersection passing behavior, due to the characteristics of straight and left turn vehicles sharing one passing interval and no special left turn signal, is one of the most complex automatic driving scenes. The interaction between left turn vehicles and opposite straight vehicles during the crossing process brings difficulties to decision and planning.

[0003] In the prior art, the unprotected intersection left turn is usually adopted by a rule-based decision-making method. Although the rule-based decision-making method is simple in logic, easy to work, and can directly draw on human driving experience, it usually ignores the complexity of traffic and is usually conservative and weak in generalization. SUMMARY

[0004] In order to solve the above technical problems, the application provides an automatic driving unprotected left turn game decision-making method and system considering interaction style, which not only guarantees interaction safety, but also improves passing efficiency.

[0005] To achieve the above purpose, the application adopts the following technical scheme:

[0006] The automatic driving unprotected left turn game decision-making method considering interaction style comprises the following steps:

[0007] The time interaction feature and the dynamic interaction feature of the target vehicle when left turning with other vehicles are obtained, and the time interaction feature and the dynamic interaction feature are normalized;

[0008] The time interaction parameter is calculated by using the normalized time interaction feature, and the dynamic interaction parameter is calculated by using the normalized dynamic interaction feature. Then, the comprehensive interaction parameter is calculated by using the time interaction parameter and the dynamic interaction parameter;

[0009] The distribution of interaction style is modeled as a mixed Gaussian distribution. Based on Bayesian inference, the confidence of the interaction style of other vehicles is updated in real time through the comprehensive interaction parameter;

[0010] A social level-k game model is established. The confidence of the updated interaction style of other vehicles is used as a weight to calculate the weighted sum of expected rewards under different interaction styles, so as to obtain a comprehensive reward function;

[0011] Based on a predefined discrete action space, the state trajectory of the vehicle in a future time domain is predicted by using a vehicle kinematics model, a Monte Carlo tree search algorithm is adopted to solve the problem by taking the comprehensive reward function as the optimization target, and the optimal decision action sequence of the target vehicle is output.

[0012] The application also provides an automatic driving unprotected left turn game decision method considering interaction style, which comprises a data acquisition module, a calculation module, an updating module, a model establishment module and a solving module.

[0013] The data acquisition module is used for acquiring the time interaction feature and the dynamic interaction feature of the target vehicle when making a left turn, and normalizing the time interaction feature and the dynamic interaction feature.

[0014] The calculation module is used for calculating the time interaction parameter by using the normalized time interaction feature and the dynamic interaction parameter by using the normalized dynamic interaction feature, and calculating the comprehensive interaction parameter by using the time interaction parameter and the dynamic interaction parameter.

[0015] The updating module is used for modeling the distribution of the interaction style as a mixed Gaussian distribution, and updating the confidence of the interaction style of the other vehicle in real time based on Bayesian inference through the comprehensive interaction parameter.

[0016] The model establishment module is used for establishing a social level-k game model, taking the confidence of the updated interaction style of the other vehicle as a weight, calculating the weighted sum of the expected reward under different interaction styles, and obtaining a comprehensive reward function.

[0017] The solving module is used for predicting the state trajectory of the vehicle in a future time domain based on a predefined discrete action space by using a vehicle kinematics model, adopting a Monte Carlo tree search algorithm fused with comfort constraints, taking the comprehensive reward function as the optimization target to solve, and outputting the optimal decision action sequence of the target vehicle.

[0018] The effects provided in the summary are only the effects of the embodiments, not all the effects of the application, and one of the above technical solutions has the following advantages or beneficial effects:

[0019] The application provides an automatic driving unprotected left turn game decision method and system considering interaction style, and belongs to the technical field of vehicle driving decision.

[0020] The time interaction parameter is calculated by using the normalized time interaction feature, and the dynamic interaction parameter is calculated by using the normalized dynamic interaction feature; the comprehensive interaction parameter is calculated by using the time interaction parameter and the dynamic interaction parameter; the distribution of the interaction style is modeled as a mixed Gaussian distribution, and the confidence of the he-vehicle interaction style is updated in real time through the comprehensive interaction parameter based on Bayesian inference; a social level-k game model is established, the confidence of the updated he-vehicle interaction style is taken as a weight, and the weighted sum of the expected rewards under different interaction styles is calculated to obtain a comprehensive reward function; based on a predefined discrete action space, the state trajectory of the vehicle in the future time domain is predicted by using a vehicle kinematics model, and a Monte Carlo tree search algorithm combined with comfort constraints is adopted to solve the optimization target of the comprehensive reward function, and the optimal decision action sequence of the target vehicle is output. The automatic driving unprotected left turn game decision method considering the interaction style is also proposed, and the automatic driving unprotected left turn game decision system considering the interaction style is also proposed. The automatic driving unprotected left turn game decision method considering the interaction style of the present application uses a level-k game model to model the interactive game relationship between the automatic driving vehicle and the human driving vehicle, and uses the proposed interaction style recognition module to recognize the distribution of the interaction style of the surrounding vehicles, thereby enhancing the sociality and interaction of the game model.

[0021] The present application can accurately estimate the interaction style of surrounding traffic participants through real-time behavior observation, so that the automatic driving vehicle exhibits human-like decision logic in complex intersection scenarios, ensuring interaction safety and improving traffic efficiency. Compared with the prior art, the present application can achieve higher traffic success rate, longer average traffic distance, higher average traffic speed and smaller average acceleration change rate. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 The schematic diagram of the automatic driving unprotected left turn game decision method considering the interaction style of the present application embodiment 1 is shown in the figure.

[0023] Figure 2 The schematic diagram of the automatic driving unprotected left turn game decision method considering the interaction style of the present application embodiment 1 is shown in the figure.

[0024] Figure 3 The schematic diagram of the automatic driving unprotected left turn game decision system considering the interaction style of the present application embodiment 2 is shown in the figure. DETAILED DESCRIPTION

[0025] For the purpose of clearly illustrating the technical features of the present application, the present application will be described in detail below with specific embodiments and in conjunction with the accompanying drawings. The disclosure below provides many different embodiments or examples for implementing the different structures of the present application. In order to simplify the disclosure of the present application, the components and arrangements of specific examples are described below. In addition, reference numerals and / or letters can be repeated in different examples. Such repetition is for the purpose of simplification and clarity, and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. It should be noted that the components illustrated in the drawings are not necessarily drawn to scale. The present application omits the description of well-known components and processing techniques and processes to avoid unnecessarily limiting the present application.

[0026] Embodiment 1

[0027] The embodiment 1 of the present application proposes an automatic driving unprotected left turn game decision method considering interaction style, which is used to solve the technical problems of ignoring the complexity of traffic, being too conservative and weak generalization in the prior art. Figure 1 The schematic diagram for realizing the automatic driving unprotected left turn game decision method considering interaction style of the embodiment 1 of the present application;

[0028] Real-time state information of the ego vehicle and the opposite interaction vehicle is obtained, and the state information includes position, speed, acceleration and heading angle;

[0029] The time interaction feature and the dynamic interaction feature of the target vehicle when turning left are selected, and the time interaction feature and the dynamic interaction feature are normalized;

[0030] The time interaction feature includes collision time and collision time rate of change; the dynamic interaction feature includes the speed of the other vehicle and the acceleration of the other vehicle.

[0031] In order to ensure the consistency of the interaction feature scale, the time interaction feature and the dynamic interaction feature are normalized, specifically:

[0032] ;

[0033] ;

[0034] ;

[0035] ;

[0036] wherein, is the normalized collision time; is the normalized collision time rate of change; is the normalized speed of the other vehicle; is the normalized acceleration of the other vehicle; is a parameter for determining the steepness of the collision time curve; The parameter that determines the steepness of the rate of change of the collision time; Parameters used to determine the steepness of the slope for other vehicles; The parameters that determine the steepness of the vehicle's acceleration are: the larger the value of the parameter that determines the steepness, the more sensitive the response. This represents the offset of the collision time. This is the offset of the rate of change of the collision time; The offset of the other vehicle's speed; The offset of the acceleration of the other vehicle.

[0037] The parameter values ​​and offsets that determine the steepness can be determined by analyzing the distribution of data from a real dataset. This can be done by observing the extracted velocity and acceleration data. and The smaller the value, the stronger the tendency to yield. and The larger the value, the stronger the tendency to yield.

[0038] To fully consider the combined effect of each interaction feature, a temporal interaction parameter (TIP) defined by the temporal interaction feature and a dynamic interaction parameter (DIP) defined by the dynamic interaction feature were designed, as follows:

[0039] ;

[0040] ;

[0041] in, For time-interaction parameters; For dynamic interactive parameters;

[0042] ;

[0043] in, For comprehensive interactive parameters.

[0044] Interaction parameters right and Take the geometric mean to balance the effects of both and ensure the result is within [0,1]. The closer the value is to 1, the stronger the tendency to yield; the closer it is to 0, the stronger the tendency to cut in front.

[0045] In complex scenarios like unsignalized intersections, the social interaction styles of other traffic participants are constantly changing. Ignoring the uncertainty of these interaction styles can lead to inappropriate or overly conservative decisions. Therefore, this invention proposes a social level-k game model that utilizes comprehensive interaction parameters. The interaction styles of interactive objects are estimated, including conservative, natural, and aggressive styles. Bayesian inference is used to update the confidence scores for each interaction style in real time. The confidence scores for aggressive, conservative, and natural styles are defined as follows: Level-0 is defined as an aggressive interaction style, level-1 is defined as a conservative interaction style, and level-2 is defined as an interaction style with an aggressiveness level between level 0 and level 1, i.e., a natural interaction style.

[0046] Then, the confidence scores of different interaction styles are used to construct a level-k game model. Specifically, the distribution of interaction styles is modeled as a Gaussian mixture distribution, and based on Bayesian inference, the confidence scores of other vehicles' interaction styles are updated in real time using the comprehensive interaction parameters.

[0047] The distribution of interaction styles is modeled as a Gaussian mixture distribution. The Gaussian mixture distribution is used to address the inherent uncertainty and complexity associated with the multimodal interaction styles of interactive objects. It expresses uncertainty with multiple modes by integrating multiple continuous probability distributions.

[0048] The mixture Gaussian distribution is represented as:

[0049] ;

[0050] in, It is a mixture Gaussian distribution function; It is the weight of each interaction style and , It is the probability density function of the distribution of each interaction style, with a mean of The covariance is .

[0051] The interaction style of interactive objects is constantly changing and needs to be updated in real time. Bayesian inference is used to determine the confidence level for these real-time updates.

[0052] ;

[0053] in, j are both interaction style indices; It is the first Posterior probability of each interaction style; It is the first The prior probability of each interaction style; It is the first The prior probability of each interaction style; It is the first The probability density function of each interaction style distribution, and in An assessment will be conducted at the site.

[0054] The social level-k game model is established as follows:

[0055] ;

[0056] ;

[0057] ; It is the first Confidence level of each interaction style; For the first One interaction style;

[0058] The social level-k model can comprehensively consider the distribution of different social interaction styles of interactive objects and can dynamically update the confidence of interaction styles.

[0059] A social level-k game model is established, and the confidence of the updated interaction style is used as the weight to calculate the weighted sum of the expected rewards under different interaction styles, so as to obtain the comprehensive reward function.

[0060] Based on a predefined discrete action space, the vehicle's state trajectory in the future time domain is predicted using a vehicle kinematics model.

[0061] The vehicle kinematics model is used to determine the current time. state vector and the selected action Calculate the next time step state vector Specifically:

[0062] ;

[0063] in, for Vehicles at all times coordinate; for Vehicles at all times coordinate; For The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; for Vehicles at all times coordinate; for Vehicles at all times coordinate; For The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; For time step;

[0064] The predefined discrete action space includes:

[0065] Hold: (0, 0); Left turn: (1, 0); Right turn: (1, 1);

[0066] Accelerate: (2.5, 0); Decelerate: (-2.5, 0); Hard brake: (-5, 0);

[0067] Fast left turn: (1, 0); Fast right turn: (1, 1).

[0068] The reward of the selection policy is defined in the form of a rolling horizon optimization. denotes the action sequence selected by the ego vehicle in the time horizon . The goal of the ego vehicle is to find the optimal action sequence that maximizes the cumulative reward and execute the first action in the action sequence . This process will be repeated at every time instant thereafter. The cumulative reward function is defined as follows:

[0069] The cumulative reward function is:

[0070] ;

[0071] ;

[0072] ;

[0073] where denotes the discount factor of the ego vehicle; has the same meaning as ;

[0074] denotes the immediate reward of the ego vehicle when it takes action under the assumption that other vehicles take the optimal action sequence ; denotes the time horizon; denotes the optimal action sequence of the ego vehicle; it indicates that the optimal action of the ego vehicle is dependent on the optimal actions of other vehicles, which reflects the interaction between agents. ​​​​​​​​​​​​

[0075] To better guide the vehicle to choose a better strategy, the immediate reward is considered as a function of multiple factors, which is defined as:

[0076] ;

[0077] The reward is divided into three types: safety reward , efficiency reward and comfort reward . , and are the weight coefficients of the three reward types respectively.

[0078] Safety reward can be expressed as:

[0079] ;

[0080] where is to avoid the vehicle to produce the behavior of violating traffic regulations, such as crossing solid lane lines, driving out of the drivable area, etc. When performing action , if the behavior of violating traffic regulations occurs, a larger penalty is returned. is defined as:

[0081] ;

[0082] is the safety distance penalty, to guide the vehicles to maintain a safe distance between them:

[0083] ;

[0084] is the vehicle distance, is the safety distance, is a constant coefficient. If the distance between vehicles is less than the safety distance, a penalty related to the difference between the vehicle distance and the safety distance is returned.

[0085] Efficiency reward to guide the agent to pass through the intersection area quickly and improve the efficiency of passing through. can be expressed as:

[0086] ;

[0087] where represents the reward to make the vehicle maintain the reference speed as much as possible, is defined as:

[0088] ;

[0089] is the reference speed, is a constant coefficient, is designed to give a greater penalty when the vehicle speed is less than 0, so should be set to a larger coefficient.

[0090] represents the reward related to the distance between the vehicle and the target position. is defined as:

[0091] ;

[0092] where is the position and heading angle of the vehicle, represents the position and heading angle of the target point.

[0093] comfort reward In order to avoid the occurrence of sharp acceleration or deceleration changes during the driving of the vehicle, the rate of change of acceleration is selected to construct :

[0094] ;

[0095] is the rate of change of acceleration, is a constant coefficient.

[0096] After constructing the level-k game model, the solution method based on Monte Carlo tree search is used to solve the game. The Monte Carlo tree search algorithm includes the following parts: selection, expansion, simulation and backtracking.

[0097] The comfort constraint is used to limit the action expansion and simulation process in the Monte Carlo tree search algorithm; the selection stage usually balances exploration and utilization through the Upper Confidence Bound (UCB) formula, that is, both good nodes that have been discovered and nodes that have not been fully explored are considered: in the expansion stage, the action that satisfies the comfort condition is expanded as a child node;

[0098] ;

[0099] where, is the average reward of the th node, is a constant for balancing exploration and utilization, is the total number of explorations, is the number of visits to this node, is the accumulated reward of this node. By calculating the UCB value of each child node and selecting the child node with the maximum value, the search process can be effectively guided.

[0100] In the simulation phase, a strategy satisfying the comfort condition is used for action selection; the comfort condition is to judge whether the acceleration rate of change exceeds the threshold.

[0101] A random strategy is usually used for action selection, which is relatively simple and can provide enough information to estimate the potential income of the node. However, the rationality of selecting actions is not considered in the random strategy, and the actions selected by the random strategy may not meet the actual requirements, which is unreasonable. Therefore, the comfort requirement must be considered in the random strategy. Similarly, in the expansion phase, only the actions that meet the comfort requirement described above are expanded into child nodes under the leaf node, rather than all actions being expanded into child nodes, is a function for judging whether it meets the comfort requirement, and the function is true if the action does not meet the comfort requirement, and vice versa:

[0102] ;

[0103] is the acceleration of the action at t+1 time, is the acceleration of the action at t time, is the limit value of jerk.

[0104] The automatic driving unprotected left turn game decision method considering interaction style provided in Embodiment 1 of the present application uses a level-k game model to model the interactive game relationship between the automatic driving vehicle and the human driving vehicle for the left turn task at the intersection without a signal, and uses the proposed interaction style recognition module to identify the distribution of the interaction style of the surrounding vehicles, thereby enhancing the sociality and interaction of the game model.

[0105] In order to fully verify the execution process of the automatic driving unprotected left turn game decision method considering interaction style provided in Embodiment 1 of the present application, Figure 2 is a schematic diagram of the intersection three-vehicle game scenario provided in Embodiment 1 of the present application; the blue vehicle is the ego vehicle, which is about to make a left turn. The red and green vehicles are the other vehicles, which are straight. Each vehicle is controlled by the method proposed in the present application. Taking the ego vehicle as an example, first, the ego vehicle calculates the collision time, the collision time rate of change, the speed of the other vehicle, and the acceleration of the other vehicle according to the sensor data. Then, according to the above data, the interaction parameters of the two other vehicles are calculated. Thereafter, according to the respective interaction parameters, the distribution of the interaction style of each other vehicle is updated through Bayesian inference. Finally, the optimal strategy of the ego vehicle is obtained by using Monte Carlo tree search. In this example, the ego vehicle can safely pass between the two other vehicles and complete the left turn.

[0106] The automatic driving unprotected left turn game decision method considering interaction style provided in embodiment 1 of the present application can accurately estimate the interaction style of surrounding traffic participants through real-time behavior observation, so that the automatic driving vehicle can exhibit human-like decision logic in complex intersection scenes, which not only ensures interaction safety, but also improves traffic efficiency, and compared with the prior art, can realize higher traffic success rate, longer average traffic distance, higher average traffic speed and smaller average acceleration change rate.

[0107] Embodiment 2

[0108] Based on the automatic driving unprotected left turn game decision method considering interaction style provided in embodiment 1 of the present application, embodiment 2 of the present application further provides an automatic driving unprotected left turn game decision system considering interaction style, Figure 3 The automatic driving unprotected left turn game decision system considering interaction style provided in embodiment 2 of the present application is shown in the accompanying drawings, and the system comprises a data acquisition module, a calculation module, an updating module, a model establishment module and a solving module.

[0109] The data acquisition module is used to acquire time interaction features and dynamic interaction features of the target vehicle when turning left and other vehicles, and normalize the time interaction features and the dynamic interaction features.

[0110] The calculation module is used to calculate time interaction parameters by using the normalized time interaction features and calculate dynamic interaction parameters by using the normalized dynamic interaction features, and then calculate comprehensive interaction parameters by using the time interaction parameters and the dynamic interaction parameters.

[0111] The updating module is used to model the distribution of interaction style as a mixed Gaussian distribution, and update the confidence of the interaction style of other vehicles in real time based on Bayesian inference through the comprehensive interaction parameters.

[0112] The model establishment module is used to establish a social level-k game model, take the confidence of the updated interaction style of other vehicles as a weight, calculate the weighted sum of expected rewards under different interaction styles, and obtain a comprehensive reward function.

[0113] The solving module is used to predict the state trajectory of the vehicle in the future time domain by using a vehicle kinematics model based on a predefined discrete action space, adopt a Monte Carlo tree search algorithm fused with comfort constraints, take the comprehensive reward function as an optimization objective to solve, and output an optimal decision action sequence of the target vehicle.

[0114] In the data acquisition module of the present application, the time interaction features include collision time and collision time change rate, and the dynamic interaction features include the speed of other vehicles and the acceleration of other vehicles.

[0115] In the calculation module, the time interaction features and the dynamic interaction features are both normalized, specifically:

[0116] ;

[0117] ;

[0118] ;

[0119] ;

[0120] wherein, is the normalized collision time; is the normalized collision time rate; is the normalized other vehicle speed; is the normalized other vehicle acceleration; is a parameter that determines the steepness of the collision time curve; is a parameter that determines the steepness of the collision time rate; is a parameter that determines the steepness of the other vehicle speed; is a parameter that determines the steepness of the other vehicle acceleration; is the offset of the collision time; is the offset of the collision time rate; is the offset of the other vehicle speed; is the offset of the other vehicle acceleration.

[0121] The time interaction parameter is calculated by using the normalized time interaction feature, and the dynamic interaction parameter is calculated by using the normalized dynamic interaction feature, and then the comprehensive interaction parameter is calculated by using the time interaction parameter and the dynamic interaction parameter, specifically as follows:

[0122] ;

[0123] ;

[0124] wherein, is the time interaction parameter; is the dynamic interaction parameter;

[0125] ;

[0126] wherein, is the comprehensive interaction parameter.

[0127] In the updating module, the Gaussian mixture distribution is represented as:

[0128] ;

[0129] wherein, is the Gaussian mixture distribution function; is the weight of each interaction style and , is the probability density function of each interaction style distribution, whose mean is and covariance is ;

[0130] The confidence of his car interaction style is updated in real time through Bayesian inference:

[0131] ;

[0132] wherein, and are interaction style indexes; is the posterior probability of the th interaction style; is the prior probability of the th interaction style; is the prior probability of the th interaction style; is the probability density function of the th interaction style distribution, and is evaluated at

[0133] The interaction style includes aggressive type, conservative type and natural type; and the interaction style corresponds to the social level-k game level, the aggressive type corresponds to level-0, the conservative type corresponds to level-1, and the natural type corresponds to level-2; the Bayesian inference is used to update the confidence of each interaction style according to the observed comprehensive interaction parameters at the current moment.

[0134] In the model establishing module, the social level-k game model is established as:

[0135] ;

[0136] ;

[0137] is the confidence of the th interaction style;

[0138] The comprehensive reward function is:

[0139] ;

[0140] ;

[0141] ;

[0142] wherein, represents the discount factor of the ego car; and have the same meaning;

[0143] This indicates that the vehicle takes the optimal sequence of actions from other vehicles. In this situation, the vehicle takes action. The instant reward received; Represents the time domain; This represents the optimal sequence of actions for the vehicle; it indicates that the optimal actions of the vehicle are obtained by relying on the optimal actions of other vehicles, which reflects the interaction between agents.

[0144] In the solution module, the vehicle kinematics model is used based on the current time. state vector and the selected action Calculate the next time step state vector Specifically:

[0145] ;

[0146] in, for Vehicles at all times coordinate; for Vehicles at all times coordinate; For The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; for Vehicles at all times coordinate; for Vehicles at all times coordinate; For The heading angle of the vehicle at any given time; for The speed of the vehicle at any given time; For time step.

[0147] The predefined discrete action space includes: (Maintaining: ( =(0,0); Turn left: ( )=( );Turn right:( )=( );accelerate:( =(2.5,0); Deceleration: ( =(-2.5,0); Emergency braking: ( = (-5, 0); Quick left turn: ( )=( Quick right turn: ( )=( )。

[0148] The Monte Carlo tree search algorithm fused with the comfort constraint is adopted to solve the optimization target of the comprehensive reward function, and the optimal decision action sequence of the target vehicle is output, specifically:

[0149] The action expansion and simulation process in the Monte Carlo tree search algorithm are restricted by the comfort constraint;

[0150] In the expansion phase, the action meeting the comfort condition is expanded as a child node;

[0151] In the simulation phase, the action selection is performed by using the strategy meeting the comfort condition; The comfort condition is to judge whether the acceleration change rate exceeds the threshold.

[0152] The automatic driving unprotected left turn game decision system considering interaction style provided in the embodiment 2 of the application, for the left turn task at the intersection without signal, uses the level-k game model to model the interactive game relationship between the automatic driving vehicle and the human driving vehicle, and uses the proposed interactive style recognition module to identify the distribution of the interactive style of the surrounding vehicles, and enhances the sociality and interactivity of the game model.

[0153] The automatic driving unprotected left turn game decision system considering interaction style provided in the embodiment 2 of the application can accurately estimate the interactive style of the surrounding traffic participants through real-time behavior observation, so that the automatic driving vehicle can exhibit human-like decision logic in complex intersection scenarios, which not only ensures the interaction safety, but also improves the traffic efficiency, and compared with the prior art, can realize higher traffic success rate, longer average traffic distance, higher average traffic speed and smaller average acceleration change rate.

[0154] The description of the related part in the automatic driving unprotected left turn game decision system considering interaction style provided in the embodiment 2 of the application can refer to the detailed description of the corresponding part in the automatic driving unprotected left turn game decision method considering interaction style provided in the embodiment 1 of the application, and will not be repeated here.

[0155] It is to be noted that, in the present text, the terms such as first and second, and the like are used merely to differentiate one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between such entities or operations. Moreover, the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements in the list. Without more limitations, an element defined by an expression "comprising a..." does not exclude the existence of additional identical elements in the process, method, article, or apparatus including the element. In addition, the above-described technical solutions provided by the embodiments of the present application have not been described in detail, which are consistent with the implementation principles of the corresponding technical solutions in the prior art, so as not to be too verbose.

[0156] The above describes the specific embodiments of the present application in conjunction with the accompanying drawings, but is not a limitation on the protection scope of the present application. Based on the above description, those skilled in the art can make other different forms of modifications or changes. Here, it is not necessary or possible to exhaust all the embodiments. Various modifications or changes made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.

Claims

1. An automatic driving unprotected left turn game decision method considering interaction style, characterized in that, The method comprises the following steps: obtaining time interaction features and dynamic interaction features of the target vehicle when turning left and a vehicle, and normalizing the time interaction features and the dynamic interaction features; calculating time interaction parameters by using the normalized time interaction features and calculating dynamic interaction parameters by using the normalized dynamic interaction features, and then calculating comprehensive interaction parameters by using the time interaction parameters and the dynamic interaction parameters; modeling a distribution of interaction styles as a mixed Gaussian distribution, updating a confidence of the interaction style of the vehicle in real time through the comprehensive interaction parameters based on Bayesian inference; establishing a social level-k game model, taking the confidence of the updated interaction style of the vehicle as a weight, calculating a weighted sum of expected rewards under different interaction styles to obtain a comprehensive reward function; based on a predefined discrete action space, predicting a state trajectory of the vehicle in a future time domain by using a vehicle kinematics model, and solving an optimization target by using the comprehensive reward function through a Monte Carlo tree search algorithm fused with comfort constraints, and outputting an optimal decision action sequence of the target vehicle.

2. The automatic driving left-turn game decision-making method considering interaction style according to claim 1, characterized in that, The time interaction features include a collision time and a collision time change rate, and the dynamic interaction features include a speed of the vehicle and an acceleration of the vehicle. 3.The method of claim 2, wherein, The time interaction features and the dynamic interaction features are normalized, and the normalization comprises: ; ; ; ; wherein, is the normalized collision time; is the normalized collision time rate of change; is the normalized other vehicle speed; is the normalized other vehicle acceleration; is a parameter that determines how steep the collision time curve is; is a parameter that determines how steep the collision time rate of change is; is a parameter that determines how steep the other vehicle speed is; is a parameter that determines how steep the other vehicle acceleration is; is an offset for the collision time; is an offset for the collision time rate of change; is an offset for the other vehicle speed; is an offset for the other vehicle acceleration.

4. The automatic driving left-turn game decision method considering interaction style according to claim 3, characterized in that, calculating time interaction parameters by using the normalized time interaction features and calculating dynamic interaction parameters by using the normalized dynamic interaction features, and then calculating comprehensive interaction parameters by using the time interaction parameters and the dynamic interaction parameters. ; ; wherein is a time interaction parameter; is a dynamic interaction parameter; ; wherein, is the integrated interaction parameter.

5. The automatic driving left-turn game decision method considering interaction style according to claim 4, characterized in that, The distribution of the interaction styles is modeled as a mixed Gaussian distribution, and a confidence of the interaction style of the vehicle is updated in real time through the comprehensive interaction parameters based on Bayesian inference. The mixed Gaussian distribution is represented as: ; wherein, is a Gaussian mixture distribution function; is a weight of each interaction style and , is a probability density function of each interaction style distribution, with mean and covariance ; The confidence of the interaction style of the vehicle is updated in real time through Bayesian inference. ; wherein, j are interaction style indices; is the posterior probability of the jth interaction style; is the prior probability of the jth interaction style; is the prior probability of the jth interaction style; is the prior probability of the jth interaction style distribution, and is evaluated at .

6. The automatic driving left-turn game decision method considering interaction style according to claim 5, characterized in that, The interaction styles include an aggressive type, a conservative type and a natural type, and the interaction styles correspond to social level-k game levels, the aggressive type corresponds to level-0, the conservative type corresponds to level-1, and the natural type corresponds to level-2, and the Bayesian inference is used to update the confidence of each interaction style according to the comprehensive interaction parameters observed at the current moment.

7. The automatic driving left-turn game decision method considering interaction style according to claim 6, characterized in that, The social level-k game model is established as: ; ; ; It is the first Confidence level of each interaction style; For the first One interaction style; The comprehensive reward function is: ; ; ; wherein, represents a discount factor of the ego vehicle; has the same meaning as has the same meaning as represents the self car taking action under the condition that other cars take optimal action sequences resulting in an immediate reward; represents the optimal action sequence of the self car; represents the time horizon.

8. The automatic driving left-turn game decision-making method considering interaction style according to claim 1, characterized in that, The vehicle kinematics model is used to determine the current time. state vector and the selected action Calculate the next time step state vector Specifically: ; wherein is the coordinate of the vehicle at the time instant ; is the coordinate of the vehicle at the time instant ; is the heading angle of the vehicle at the time instant ; is the speed of the vehicle at the time instant ; is the coordinate of the vehicle at the time instant ; is the coordinate of the vehicle at the time instant ; is the heading angle of the vehicle at the time instant ; is the time step The predefined discrete action space includes: (Maintaining: ( = (0,0); Turn left: ( )=( );Turn right:( )=( );accelerate:( =(2.5,0); Deceleration: ( =(-2.5,0); Emergency braking: ( = (-5, 0); Quick left turn: ( )=( Quick right turn: ( )=( ).

9. The automatic driving left-turn game decision-making method considering interaction style according to claim 1, characterized in that, The Monte Carlo tree search algorithm fused with the comfort constraints is used to solve the optimization target, and an optimal decision action sequence of the target vehicle is outputted, and the solving comprises: the action expansion and the simulation process in the Monte Carlo tree search algorithm are limited by the comfort constraints; in the expansion stage, an action satisfying a comfort condition is expanded as a child node; in the simulation stage, an action is selected by using a strategy satisfying the comfort condition; and the comfort condition is that whether an acceleration change rate exceeds a threshold value.

10. An automatic driving unprotected left turn game decision method considering interaction style, characterized in that, The method comprises a data acquisition module, a calculation module, an updating module, a model establishment module and a solving module. The data acquisition module is used to obtain time interaction features and dynamic interaction features of the target vehicle when turning left and a vehicle, and to normalize the time interaction features and the dynamic interaction features. The computing module is configured to calculate a time interaction parameter by using the normalized time interaction feature and a dynamic interaction parameter by using the normalized dynamic interaction feature; and calculate a comprehensive interaction parameter by using the time interaction parameter and the dynamic interaction parameter; The updating module is configured to model a distribution of the interaction style as a Gaussian mixture distribution, and update a confidence of the interaction style of the other vehicle in real time based on Bayesian inference and the comprehensive interaction parameter; The model establishing module is configured to establish a social level-k game model, take the confidence of the updated interaction style of the other vehicle as a weight, calculate a weighted sum of expected rewards under different interaction styles, and obtain a comprehensive reward function; The solving module is configured to predict a state trajectory of the vehicle in a future time domain by using a vehicle kinematics model based on a predefined discrete action space, adopt a Monte Carlo tree search algorithm fused with a comfort constraint, take the comprehensive reward function as an optimization objective, and output an optimal decision action sequence of the target vehicle.

Citation Information

Patent Citations

  • Intelligent vehicle behavior decision-making method, planning method, system and storage medium

    CN114919578A

  • Game theoretic decision making

    US20230182014A1