Lane changing decision-making method based on k-level mixed strategy game
By establishing a driver model for interactive vehicles and using k-level hybrid strategy game, lane-changing decisions are optimized, solving the problem that existing technologies fail to consider the characteristics of surrounding vehicles and improving the rationality and safety of lane-changing decisions.
Patent Information
- Application Number
- CN202511078586.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-14
AI Technical Summary
Existing lane-changing decision-making methods based on game theory fail to fully consider the driving characteristics and deeper thinking of surrounding vehicles, which may lead to misjudgments and safety hazards.
A driver model for interactive vehicles is established. The vehicle's trajectory under different strategies is predicted through k-level hybrid strategy game. Driver style recognition is performed by combining fuzzy logic rules to optimize lane-changing decisions and improve rationality and safety.
It enhances the understanding of the surrounding traffic environment, simulates the deep thinking process of vehicle interaction, reduces the occurrence of safety accidents, and ensures safety and efficiency when changing lanes.
Smart Images

Figure CN120954261A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle control technology, and specifically to a lane-changing decision-making method based on k-level hybrid strategy game. Background Technology
[0002] Autonomous driving technology, with its potential to reduce traffic accidents, improve traffic efficiency, and ensure passenger comfort, has become a hot topic in the automotive industry. The application of autonomous driving mobility services is increasingly welcomed by mobility service providers and technology companies. When intelligent vehicles drive in structured scenarios such as cities and highways, their driving behavior can be divided into lane keeping and lane changing. Autonomous driving requires intelligent vehicles to make real-time decisions regarding driving behavior.
[0003] Currently, game-theoretic lane-changing decision-making methods mainly establish strategy sets for the controlled vehicle and the target lane vehicle, and select the optimal strategy as the lane-changing decision for the controlled vehicle. For example, invention patent with patent number "CN202110608331.3" provides a game-theoretic-based lane-changing decision-making method, including the following steps: modeling the lane-changing scenario in a game, setting the participants in the game process as the vehicle to be changed lanes and the target lane vehicle; setting the strategy set for the vehicle to be changed lanes and the strategy set for the target lane vehicle; calculating all possible strategy combinations during the lane-changing process; calculating the game payoffs of the vehicle to be changed lanes and the target lane vehicle after a successful lane change based on the strategy combinations, and calculating the final payoffs of the vehicle to be changed lanes and the target lane vehicle; constructing a joint payoff matrix based on the final payoffs of the vehicle to be changed lanes and the target lane vehicle; calculating the expected total payoffs of the vehicle to be changed lanes and the target lane vehicle based on the joint payoff matrix; and calculating the mixed strategy and expected payoff of the vehicle to be changed lanes and the target lane vehicle at the Nash equilibrium state during the lane-changing game process based on the expected total payoffs of the vehicle to be changed lanes and the target lane vehicle. However, the patent does not take into account the driving characteristics of surrounding vehicles that may interact with the controlled vehicle, nor does it consider the different strategies that these interacting vehicles may adopt under deeper consideration, and the impact of their driving trajectories on the strategies that the controlled vehicle will adopt. Therefore, when making lane-changing decisions, it may misjudge the future driving situation of the interacting vehicles, thereby causing safety hazards. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a lane-changing decision-making method based on k-level hybrid strategy game theory. By establishing a driver model of the interacting vehicle, it enhances the in-depth understanding of the surrounding traffic environment. Furthermore, it can simulate the deep thinking process of the controlled vehicle when interacting with surrounding vehicles, improving the rationality of lane-changing decisions and thus effectively reducing the occurrence of safety accidents.
[0005] The technical solution adopted in this invention is as follows:
[0006] An embodiment of the present invention proposes a lane-changing decision-making method based on k-level hybrid strategy game, comprising the following steps: S1, establishing a driver model of the interactive vehicles, the interactive vehicles including the front and rear vehicles located on the front and rear sides of the controlled vehicle; S2, through k-level hybrid strategy game, based on the environmental information of the controlled vehicle and the driver model of the interactive vehicles, predicting the driving trajectories of the controlled vehicle and the interactive vehicles under different strategies, obtaining the expected payoffs of the strategies chosen by the controlled vehicle and the rear vehicle in the target lane at different levels of inference depth, the strategies of the controlled vehicle including lane changing and not changing lanes, and the strategies of the rear vehicle in the target lane including yielding and not yielding; S3, solving for the Pareto optimal objective of maximizing the sum of the expected payoffs of the controlled vehicle and the rear vehicle in the target lane, to obtain the lane-changing decision of the controlled vehicle.
[0007] In addition, the lane-changing decision-making method proposed according to the present invention may also have the following additional technical features:
[0008] According to an embodiment of the present invention, step S1 specifically includes: S11, based on the historical driving trajectories of the following vehicles and their preceding and following vehicles in the interactive vehicles, a genetic algorithm is used to solve for the driver model of the following vehicles; S12, based on the historical driving trajectory of the natural vehicles in the preceding vehicles that do not have following behavior, the driver model of the natural vehicles is established as a uniform speed model.
[0009] According to one embodiment of the present invention, in step S2, the driving trajectory of the controlled vehicle when it does not change lanes is obtained by following the preceding vehicle based on a preset driver model of the controlled vehicle; the driving trajectory of the controlled vehicle when it changes lanes is obtained based on the trapezoidal acceleration method and the predicted driving trajectory of the interacting vehicle.
[0010] According to one embodiment of the present invention, in step S2, the strategy of a vehicle at inference depth k is obtained by considering the probability of other vehicles at inference depth k-1 choosing different strategies. The probability of a vehicle choosing different strategies at inference depth 0 is the same, and k is an integer greater than or equal to 1.
[0011] According to one embodiment of the present invention, the driver model of the following vehicle is:
[0012]
[0013] Where a represents acceleration, a max v is the maximum acceleration that the following vehicle can achieve without restrictions, and v is the current following speed of the following vehicle. desUnder ideal conditions, without any other vehicles obstructing the view, the maximum speed that the driver of the following vehicle expects to reach is δ, where δ is the acceleration exponent, representing the driver's reaction intensity to the difference between the following speed and the expected speed; s * (v, Δv) represents the desired distance between the following vehicle and the vehicle in front, Δv represents the speed difference between the following vehicle and the vehicle in front, s0 refers to the minimum distance maintained between the following vehicle and the vehicle in front when stationary, T is the minimum time interval that the driver expects to maintain while following, and b is the deceleration speed that the driver feels comfortable with during deceleration. This is used to ensure that the desired spacing is not negative.
[0014] According to one embodiment of the present invention, before step S2, the method further includes: based on the driver model of the interactive vehicle, using fuzzy logic rules to perform driver style recognition on the interactive vehicle, so as to take the driver style of the vehicle into account when calculating the expected benefits of the following vehicle.
[0015] According to one embodiment of the present invention, when performing driver style recognition on the interactive vehicle using fuzzy logic rules, the maximum acceleration parameter 'a' in the driver model of the interactive vehicle is used as the basis. max The initial driver style of the interactive vehicle is obtained by taking into account the following distance parameter T. Then, based on the initial driver style of the interactive vehicle, the observation duration observeT when the driver model was established is considered to obtain the final driver style of the interactive vehicle.
[0016] According to one embodiment of the present invention, the expected benefits of the controlled vehicle include speed benefits, safety benefits, and following distance benefits, and the expected benefits of the vehicle behind in the target lane include speed benefits and following distance benefits.
[0017] According to one embodiment of the present invention, when the controlled vehicle has lane-changing space on both its left and right sides, in step S2, two k-level mixed strategy games are played respectively based on the controlled vehicle and the vehicle behind it in its left lane, and the controlled vehicle and the vehicle behind it in its right lane. In step S3, the lane-changing decisions of the controlled vehicle corresponding to these two k-level mixed strategy games are obtained respectively. Then, the total expected payoffs of the controlled vehicle and the vehicles behind it in the left and right lanes are compared under the two lane-changing decisions respectively, and the lane-changing decision corresponding to the lane with the highest total expected payoff is selected as the final lane-changing decision of the controlled vehicle.
[0018] According to an embodiment of the present invention, in step S3, when the lane-changing decision of the controlled vehicle is to change lanes, and the strategy of the vehicle behind in the target lane under the lane-changing strategy is not to yield, the total expected benefits of the lane-changing and yielding strategy combination and the non-lane-changing and non-yielding strategy combination are compared, and the lane-changing decision of the controlled vehicle is obtained according to the strategy combination with the higher total expected benefits.
[0019] The beneficial effects of this invention are:
[0020] 1. The lane-changing decision-making method based on k-level hybrid strategy game of the present invention enhances the in-depth understanding of the surrounding traffic environment by establishing a driver model of the interacting vehicles; through k-level hybrid strategy game, based on the environmental information of the controlled vehicle and the driver model of the interacting vehicles, it predicts the driving trajectories of the controlled vehicle and the interacting vehicles under different strategies, obtains the expected payoffs of the strategies chosen by the controlled vehicle and the following vehicle in the target lane at different levels of inference depth, and solves the Pareto optimal objective by maximizing the sum of the expected payoffs of the controlled vehicle and the following vehicle in the target lane, thus obtaining the lane-changing decision of the controlled vehicle. This simulates the deep thinking process of the controlled vehicle when interacting with surrounding vehicles, improves the rationality of lane-changing decisions, and thus effectively reduces the occurrence of safety accidents.
[0021] 2. The lane-changing decision method of the present invention further plans the lane-changing trajectory of the controlled vehicle when changing lanes based on trapezoidal acceleration programming. Since the predicted driving trajectory of the interacting vehicles is fully considered, the optimality of lane-changing time and longitudinal lane-changing distance in terms of safety and efficiency is ensured.
[0022] 3. The lane-changing decision method of the present invention further combines fuzzy logic rules for driver style recognition and takes into account the observation time when building the driver model, thereby improving the accuracy of vehicle driving behavior analysis.
[0023] 4. The lane-changing decision-making method of the present invention further considers speed benefits, following distance benefits, safety benefits and driver style to make decisions, thereby ensuring the optimality of the target lane and lane-changing timing.
[0024] 5. The lane-changing decision-making method of the present invention further proposes a lane-changing decision-making method under an unsafe strategy combination, which obtains the lane-changing decision of the controlled vehicle as the lane-changing decision and the strategy of the vehicle behind the target lane under the lane-changing strategy. This ensures decision-making safety and improves the decision-making ability in complex working conditions. Attached Figure Description
[0025] Figure 1 This is a flowchart of the lane-changing decision method according to an embodiment of the present invention;
[0026] Figure 2 This is a schematic diagram of a driver identification model based on a genetic algorithm.
[0027] Figure 3 A schematic diagram of the driving trajectories of the controlled vehicle and the interacting vehicle;
[0028] Figure 4 A strategy diagram for the controlled vehicle and the vehicle behind it in the target lane:
[0029] Figure 5 Trapezoidal acceleration curves with different peak values;
[0030] Figure 6 The trapezoidal acceleration curve corresponding to the optimal lane change time;
[0031] Figure 7 A diagram illustrating a controlled vehicle changing lanes and a subsequent vehicle in the target lane not yielding;
[0032] Figure 8 A schematic diagram of a multi-vehicle lane-changing game involving controlled vehicles and interacting vehicles. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] like Figure 1 As shown in the figure, the lane-changing decision-making method based on k-level hybrid strategy game of the present invention includes the following steps:
[0035] S1. Establish a driver model for the interactive vehicle, which includes the front and rear vehicles located on the front and rear sides of the controlled vehicle.
[0036] In one embodiment of the present invention, step S1 specifically includes the following steps S11 to S12:
[0037] S11. Based on the historical driving trajectories of the following vehicles and the vehicles in front and behind them in the interactive vehicles, a genetic algorithm is used to obtain the driver model of the following vehicles.
[0038] Specifically, the driver model for the following vehicle is:
[0039]
[0040] Where a represents acceleration, a max v is the maximum acceleration that the following vehicle can achieve without restrictions, and v is the current following speed of the following vehicle. des Under ideal conditions, without any other vehicles obstructing the view, the maximum speed that the driver of the following vehicle expects to reach is δ, where δ is the acceleration exponent, representing the driver's reaction intensity to the difference between the following speed and the expected speed; s *(v, Δv) represents the desired distance between the following vehicle and the vehicle in front, Δv represents the speed difference between the following vehicle and the vehicle in front, s0 refers to the minimum distance that the following vehicle should maintain when stationary, T is the minimum time interval that the driver expects to maintain while following, and b is the deceleration that the driver feels comfortable with during deceleration, usually referring to the deceleration to avoid emergency braking. This is used to ensure that the desired spacing is not negative.
[0041] Due to the driver model [a max ,v des The six parameters [δ, s0, T, b] are all related to the car-following behavior characteristics. For a vehicle exhibiting car-following behavior, the speeds and positions of the following vehicle and the preceding vehicle are recorded at multiple times, and the six parameters are solved to obtain the driver model of the car-following vehicle. Using the recorded trajectories of the preceding and following vehicles as input, a genetic algorithm can be used to solve for the driver model of the car-following vehicle. This problem can be expressed as:
[0042] [a max ,v des ,δ,s0,T,b]=GA(TraFront,TraBack) (3)
[0043] like Figure 2 As shown, the process of driver model identification using a genetic algorithm can be divided into seven steps. Step 1: Collect the trajectories [TraFront, TraBack] of vehicles exhibiting car-following behavior and the vehicle in front. The trajectory information can include the vehicle's speed and position at each moment. Step 2: Randomly generate a set of solutions called a population. Each solution represents a potential driver model, usually represented using binary encoding. Step 3: For each solution in the population, evaluate its fitness using a fitness function based on the performance of the driver model it represents. The higher the fitness, the greater the probability of the individual being selected. Step 4: Determine if the algorithm meets the termination conditions, such as reaching the maximum number of iterations or the fitness reaching a preset threshold. If the termination conditions are met, stop iterating and output the individual with the highest fitness in the current population as the final solution; otherwise, continue iterating. Step 5: Based on the fitness of individuals, select individuals from the current population to serve as parents for the next generation; Step 6: Through crossover, parent individuals exchange certain characteristics to explore new solutions in the solution space, increasing population diversity; Step 7: Perform mutation operations on individuals, randomly changing certain characteristics to help the algorithm escape local optima and increase population diversity. Repeat the above seven steps to gradually optimize the parameters of the driver model [a] max ,v des The fitness function in the genetic algorithm is F(v, δ, s0, T, b], until the optimal model that meets the requirements is found. sim ,v,x sim x) can be:
[0044]
[0045] in, and v represents the velocity and position of the potential driver model corresponding to a single solution at time t. t and x t The values represent the actual speed and position of the following vehicle at time t, w1 and w2 are the weights of the error term, and n represents the maximum number of iterations.
[0046] S12, based on the historical driving trajectory of the natural vehicle that did not follow the car in front, the driver model of the natural vehicle is established as a uniform speed model, that is, the speed of the natural vehicle in the next moment is the same as the speed in the current moment.
[0047] by Figure 3 As shown in one specific embodiment, the vehicle in the center of the diagram is the controlled vehicle. The vehicles veh1, veh2, veh3, veh4, veh5, and veh6 surrounding the controlled vehicle are interacting vehicles. Since vehicles veh1, veh2, and veh3 exhibit car-following behavior, a genetic algorithm is used to solve for the driver models of these car-following vehicles, and trajectory prediction is performed based on their driver models. Since vehicles veh4, veh5, and veh6 do not exhibit car-following behavior and are considered to be driving naturally, a uniform speed model is used to predict their future trajectories.
[0048] S2, through k-level mixed-strategy game theory, predicts the trajectories of the controlled vehicle and the interacting vehicle under different strategies based on environmental information and driver models of the interacting vehicles. It obtains the expected payoffs of the strategies chosen by the controlled vehicle and the vehicle following in the target lane at different levels of inference depth. The controlled vehicle's strategies include lane changing and not changing lanes, and the target lane's strategies include yielding and not yielding. Figure 4 As shown in the figure. EV represents the controlled vehicle, R represents the vehicle behind in the target lane, F represents the vehicle in front in the target lane, eF represents the vehicle in front of the controlled vehicle before the lane change, LC represents the lane-changing strategy, NLC represents the no-lane-changing strategy, Y represents the avoidance strategy, and NY represents the no-avoidance strategy.
[0049] In one embodiment of the present invention, in step S2, the driving trajectory of the controlled vehicle when it does not change lanes is obtained by following the preceding vehicle based on a preset driver model of the controlled vehicle; the driving trajectory of the controlled vehicle when it changes lanes is obtained based on the trapezoidal acceleration method and the predicted driving trajectory of the interacting vehicle.
[0050] The lane-changing trajectory of a controlled vehicle during a lane change needs to consider the relationship between its longitudinal position and time. The unknowns include: the end time of the lane change, its longitudinal position, and its speed at the end time. For example... Figure 5As shown, a trapezoidal acceleration curve is used to solve for the lane-changing trajectory. When the planned peak acceleration is less than the peak acceleration for comfort, the acceleration change is represented by the solid curve in the figure; otherwise, it changes along the dashed curve. Assuming the initial acceleration of the lane change is 0, the final acceleration is 0, and the final speed is the speed of the vehicle in front in the target lane, if the lane-changing time is determined, the lane-changing trajectory can be calculated using the trapezoidal acceleration curve.
[0051] like Figure 6 As shown, in one specific embodiment, the lane-changing time is sampled from 2s at 0.2s intervals to 8s. After removing curves that do not meet the maximum longitudinal acceleration, the acceleration curves corresponding to the initial lane-changing trajectories under different lane-changing times can be obtained. Based on the predicted driving trajectory of the interactive vehicle and the lane-changing trajectories sampled in the figure, the cost of each trajectory is calculated by comprehensively considering lane-changing time, maximum longitudinal acceleration, and collision safety. The lane-changing trajectory with the minimum cost, i.e., the optimal solution, is selected as the lane-changing strategy for the controlled vehicle. Considering lane-changing efficiency, the lane-changing time is desired to be as short as possible; considering lane-changing comfort, the acceleration is desired to be as low as possible.
[0052] Safety considerations for lane changes, such as Figure 7 As shown, since the controlled vehicle faces a relatively high risk of collision with the preceding vehicle eF and the following vehicle R in the target lane at the midpoint of the lane change, and a relatively high risk of collision with the preceding vehicle F in the target lane at the end of the lane change, the time distance between the two vehicles at these moments is considered as a safety factor. The final cost function formula is:
[0053]
[0054] Among them, w t w j w eF w F and w R Represents the time weight, acceleration weight, safe distance weight to the vehicle in front in the controlled vehicle's lane, safe distance weight to the vehicle in front in the target lane, and safe distance weight to the vehicle behind in the target lane, t f This represents the lane change time, where w t w j w eF w F and w R The value can be an empirical value or adjusted according to actual needs. Represents 0.5t f The time distance between the controlled vehicle and the vehicle ahead (eF) at all times. Represents t f The time distance between the controlled vehicle and the vehicle F in front in the target lane at all times. Represents 0.5t f The time distance between the controlled vehicle and the vehicle R behind in the target lane at all times.
[0055] In one embodiment of the present invention, the strategy of a vehicle at inference depth k can be obtained by considering the probability of other vehicles at inference depth k-1 choosing different strategies. The probability of a vehicle choosing different strategies at inference depth 0 is the same, where k is an integer greater than or equal to 1.
[0056] For vehicle j at reasoning depth k, the optimal policy π* is defined by the following equation:
[0057]
[0058] Among them, S j It is the set of situational states that the driver of vehicle j can choose from. Let S represent the set of optimal policies for all vehicles with a reasoning depth of k-1, excluding vehicle j. E represents the set of optimal policies for all possible states S. j Strategy π j and The expected return value.
[0059] In one embodiment of the present invention, the expected benefits of the controlled vehicle include speed benefits, safety benefits, and following distance benefits. Since the safety benefits and following distance benefits of the vehicle behind in the target lane are equivalent, the expected benefits of the vehicle behind in the target lane include speed benefits and following distance benefits.
[0060] Speed gains from a combination of strategies: controlled vehicle lane change with following vehicles in the target lane yielding, and controlled vehicle lane change with following vehicles in the target lane not yielding. for:
[0061]
[0062] Among them, LCv(t) end ) and LCv(t0) represent the driving trajectory of the controlled vehicle under the lane-changing strategy at the lane-changing termination time t. end And the velocity at the initial time t0, The subscript indicates the strategy of the controlled vehicle, the superscript indicates the strategy of the vehicle following in the game, LC represents lane change, Y represents yielding, and NY represents not yielding.
[0063] Safety benefits under the combination of controlled vehicle lane changing and target lane rear vehicle avoidance strategies The formula is:
[0064]
[0065] FRT0=min((Fx(t0)-Yx(t0)) / Yv(t0),3) (9)
[0066] CRTE1=min((LCx(tend )-Yx(t end )) / Yv(t end ),3) (10)
[0067] Where Yx(t) and Yv(t) represent the predicted position and speed of the following vehicle under the avoidance strategy of the following vehicle in the target lane, while Fx(t) and Fv(t) represent the predicted position and speed of the preceding vehicle in the target lane at a constant speed, and LCx(t) and LCv(t) represent the predicted position and speed of the controlled vehicle under the lane-changing strategy at different times. To avoid excessive concern about following distance under safe driving conditions, a 3S following distance rule is introduced when calculating safety benefits. That is, when calculating the following distance between the controlled vehicle and the following vehicle in the target lane, formulas (1.9) and (1.10) are used, where FRT0 represents the minimum of the following distance when the controlled vehicle changes lanes and the 3S following distance; CRTE1 represents the minimum of the following distance when the following vehicle in the target lane chooses the avoidance strategy when the controlled vehicle changes lanes and the 3S following distance; t0 and t end These represent the start and end times of the controlled vehicle changing lanes, respectively.
[0068] Safety benefits under the combined strategy of controlled vehicle lane changing and subsequent vehicles in the target lane not yielding. The formula is:
[0069]
[0070] CRTE2=min((LCx(t0)-NYx(t0)) / NYv(t0),3) (12)
[0071] Where NTx(t) and NTv(t) represent the predicted position and speed of the following vehicle in the target lane at different times under the no-yield strategy, and CRTE2 represents the minimum of the following distance between the following vehicle in the target lane and the 3s following distance when the controlled vehicle changes lanes and the following vehicle chooses the no-yield strategy.
[0072] The following distance benefit of the strategy combination of controlled vehicle lane change and subsequent vehicle yielding in the target lane versus the strategy combination of controlled vehicle lane change and subsequent vehicle not yielding in the target lane. The formula is:
[0073]
[0074] eFCT0=min((eFx(t0)-LCx(t0)) / LCv(t0),3) (14)
[0075] FCTE=min((Fx(t end )-LCx(t end )) / LCv(t end ),3) (15)
[0076] Where eFx(t0) represents the predicted position of the controlled vehicle at the moment of lane change, assuming the preceding vehicle in the lane is traveling at a constant speed, and Fx(t end ) represents the predicted position of the controlled vehicle at the moment the lane change ends, assuming the vehicle in front of it in the target lane is traveling at a constant speed. eFCT0 represents the minimum following distance between the controlled vehicle and the vehicle in front in its lane at the moment the lane change begins, and the minimum following distance in 3 seconds. FCTE represents the minimum following distance between the controlled vehicle and the vehicle in front in its target lane at the moment the lane change ends, and the minimum following distance in 3 seconds.
[0077] From the above formula, we can obtain the expected return of the controlled vehicle under the combination of lane-changing avoidance strategies. Expected returns of the combination of lane-changing and non-yielding strategies They are respectively:
[0078]
[0079]
[0080] Among them, w V ,w S and w T These represent the velocity weight, safety weight, and space weight, respectively, all ranging from (0,1), and w v +w S +w T =1.
[0081] The same method can be used to obtain the expected benefits of controlled vehicles avoiding collisions without changing lanes. Expected benefits of not changing lanes or giving way for:
[0082]
[0083]
[0084] in, These represent the speed gain, safety gain, and following distance gain of the controlled vehicle under the combination of non-lane-changing avoidance strategies, respectively. These represent the speed gain, safety gain, and following distance gain of the controlled vehicle under the strategy of not changing lanes and not avoiding the vehicle, respectively. The specific calculation formulas can be obtained by referring to the above formulas (7) to (17), and will not be repeated here.
[0085] Similarly, the expected benefits of the vehicle behind in the target lane changing lanes to avoid the obstacle can be obtained by referring to the above formulas (7) to (17). The expected benefits of changing lanes without yielding Expected benefits of not changing lanes to avoid Expected benefits of not changing lanes or giving way for:
[0086]
[0087]
[0088]
[0089]
[0090] in, These represent the speed gain and following distance gain of the vehicle behind in the target lane when using a combination of lane-changing avoidance strategies, respectively. These represent the speed gain and following distance gain of the vehicle behind in the target lane when changing lanes without yielding, respectively. These represent the speed gain and following distance gain of the vehicle behind in the target lane when using the non-lane-changing avoidance strategy combination, respectively. These represent the speed gain and following distance gain for the vehicle behind in the target lane when using the strategy of not changing lanes or yielding.
[0091] In one embodiment of the present invention, before step S2, the method may further include: based on the driver model of the interactive vehicle, using fuzzy logic rules to identify the driver style of the interactive vehicle, and processing the uncertainty and complexity of the driver style of the surrounding vehicles, so as to consider the driver style of the vehicle when calculating the expected benefit of the following vehicle.
[0092] Specifically, when using fuzzy logic rules to perform driver style recognition for interactive vehicles, the driver model of the interactive vehicle can be analyzed using the 'a' parameter. max T and observeT are used as inputs, based on the maximum acceleration parameter a in the driver model of the interactive vehicle. max The initial driver style of the interactive vehicle is obtained by combining the following distance parameter T, as shown in Table 1.
[0093] Table 1
[0094]
[0095] Then, based on the initial driver style of the interactive vehicle, the observation duration observeT when building the driver model is considered to obtain the final driver style of the interactive vehicle, as shown in Table 2.
[0096] Table 2
[0097]
[0098] Expected benefits for vehicles following in the target lane, taking into account driver style They are respectively:
[0099]
[0100]
[0101]
[0102]
[0103] Where, w character This represents different driving styles of the vehicle behind in the target lane. When the driving style of the vehicle behind in the target lane is aggressive, it has high expectations for its own gains. character Take a larger value, w character The value range is (0,1], for example, w for a vehicle with a conservative driving style. character Set a specific value greater than 0 and less than 0.5 to the w value of vehicles with a conservative driving style. character Setting it to 0.5 will reduce the w of vehicles with a conservative driving style. character Let it be a specific value in (0.5, 1]. This is achieved by introducing w. character This approach addresses the impact of different driving styles of vehicles following in the target lane on their strategies during the game, thereby improving the accuracy of game interaction.
[0104] Understandably, based on the driver model and benefit calculation formula mentioned above, the driving trajectories of the controlled vehicle and the vehicle behind in the target lane can be obtained under each strategy combination. Then, the expected benefits of the controlled vehicle and the vehicle behind in the target lane under different strategy combinations can be calculated, as shown in Table 3.
[0105] Table 3
[0106]
[0107]
[0108] In a specific embodiment of the present invention, the probabilities of the controlled vehicle and the vehicle following in the target lane making different strategies in a k-level mixed strategy game are shown in Table 4 below:
[0109] Table 4
[0110]
[0111] Taking a single game in the second-order game in Table 4 as an example, assume that initially, the controlled vehicle and the vehicle behind it in the target lane have no in-depth understanding of each other, and both are initially considered level 0 players, with equal probability of adopting each strategy. In the first round of the game, both sides act as level 1 players, considering the strategy probability of the other party as a level 0 player, then recalculating their own strategy payoffs and adjusting their strategy probabilities accordingly. According to Tables 3 and 4, after considering the strategy probability of the other party as a level 0 player, the expected lane-changing payoff MG for level 1 is obtained when the controlled vehicle is a level 1 player. LC And the expected return of MG without changing lanes NLC for:
[0112]
[0113]
[0114] Among them, Q 01 Q represents the probability that a player of level 0 will avoid a vehicle following in the target lane. 02 The probability that a vehicle behind in the target lane will not yield when the player is a level 0 player.
[0115] When the controlled vehicle is played by a level 1 player, the probability of lane changing is P. 11 The probability of not changing lanes P 12 for:
[0116] P 11 =MG LC / (MG LC +MG NLC (30)
[0117] P 12 =MG NLC / (MG LC +MG NLC (31)
[0118] In the second round of the game, both players, acting as level 2 players, consider the opponent's strategy when they were level 1 players and further adjust their own strategies. Through step-by-step reasoning, they ultimately obtain the final probability of choosing each strategy and select the strategy with the highest probability from their respective strategies as the final decision strategy.
[0119] S3 is solved with the goal of maximizing the sum of the expected benefits of the controlled vehicle and the following vehicle as the Pareto optimal objective, thus obtaining the lane-changing decision of the controlled vehicle.
[0120] In some embodiments of the present invention, through k-level mixed strategy game, an unsafe situation may eventually occur where the vehicle changes lanes but does not yield. In step S3, when the controlled vehicle's lane-changing decision is to change lanes, and the strategy of the vehicle behind in the target lane under this lane-changing strategy is not to yield, the total expected payoff of the lane-changing and yielding strategy combination and the non-lane-changing and non-yielding strategy combination can be compared, and the lane-changing decision of the controlled vehicle is obtained according to the strategy combination with the higher total payoff. If If lane changing and yielding are chosen as the coordination strategy, the controlled vehicle's lane-changing decision is to change lanes; otherwise, if neither lane changing nor yielding is chosen as the coordination strategy, the controlled vehicle's lane-changing decision is not to change lanes.
[0121] like Figure 8 As shown, in one embodiment of the present invention, when the controlled vehicle has lane-changing space on both its left and right sides, in step S2, two k-level mixed strategy games are played, one between the controlled vehicle and the vehicle behind it in its left lane, and the other between the controlled vehicle and the vehicle behind it in its right lane. Specifically, the controlled vehicle establishes a k-level mixed strategy game GMAE1 with the vehicle in its left lane, and then establishes GMAE2 with the vehicle in its right lane. In step S3, the lane-changing decisions of the controlled vehicle corresponding to these two k-level mixed strategy games are obtained. Then, the total expected payoffs of the controlled vehicle and the vehicles behind it in both lanes under the two lane-changing decisions are compared, and the lane-changing decision corresponding to the lane with the highest total expected payoff is selected as the final lane-changing decision of the controlled vehicle. This allows the multi-vehicle game framework established by the present invention to not only perform lane-changing games in two-lane scenarios, but also to extend to lane-changing game decisions in four-lane and above scenarios.
[0122] The lane-changing decision-making method based on k-level hybrid strategy game according to embodiments of the present invention enhances the in-depth understanding of the surrounding traffic environment by establishing a driver model of the interacting vehicles. Through k-level hybrid strategy game, based on the environmental information of the controlled vehicle and the driver model of the interacting vehicles, the method predicts the driving trajectories of the controlled vehicle and the interacting vehicles under different strategies, obtains the expected payoffs of the strategies chosen by the controlled vehicle and the following vehicle in the target lane at different levels of inference depth, and solves the Pareto optimality objective by maximizing the sum of the expected payoffs of the controlled vehicle and the following vehicle in the target lane, thus obtaining the lane-changing decision of the controlled vehicle. This method simulates the deep thinking process of the controlled vehicle when interacting with surrounding vehicles, improves the rationality of lane-changing decisions, and effectively reduces the occurrence of safety accidents.
[0123] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. "A plurality of" means two or more, unless otherwise explicitly specified.
[0124] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0125] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0126] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0127] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0128] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0129] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A lane-changing decision-making method based on k-level mixed strategy game, characterized in that, Includes the following steps: S1, Establish a driver model for the interactive vehicle, which includes a front vehicle and a rear vehicle located on the front and rear sides of the controlled vehicle. S2, through k-level mixed strategy game, based on the environmental information of the controlled vehicle and the driver model of the interactive vehicle, predict the driving trajectory of the controlled vehicle and the interactive vehicle under different strategies, and obtain the expected payoff of the strategies chosen by the controlled vehicle and the vehicle behind the target lane at different levels of inference depth. The strategies of the controlled vehicle include changing lanes and not changing lanes, and the strategies of the vehicle behind the target lane include avoiding and not avoiding. S3, with the goal of maximizing the sum of the expected benefits of the controlled vehicle and the vehicle behind in the target lane as the Pareto optimal objective, the lane-changing decision of the controlled vehicle is obtained.
2. The lane-changing decision-making method according to claim 1, characterized in that, Step S1 specifically includes: S11, Based on the historical driving trajectories of the following vehicles and their preceding and following vehicles in the interactive vehicles, a genetic algorithm is used to obtain the driver model of the following vehicle. S12, based on the historical driving trajectory of the natural vehicle that did not follow the car in front, the driver model of the natural vehicle is established as a uniform speed model.
3. The lane-changing decision-making method according to claim 1, characterized in that, In step S2, the driving trajectory of the controlled vehicle when it does not change lanes is obtained by following the preceding vehicle based on the preset driver model of the controlled vehicle; the driving trajectory of the controlled vehicle when it changes lanes is obtained based on the trapezoidal acceleration method and the predicted driving trajectory of the interacting vehicle.
4. The lane-changing decision-making method according to claim 1, characterized in that, In step S2, the strategy of a vehicle at inference depth k is obtained by considering the probability of other vehicles at inference depth k-1 choosing different strategies. The probability of a vehicle choosing different strategies at inference depth 0 is the same, and k is an integer greater than or equal to 1.
5. The lane-changing decision-making method according to claim 2, characterized in that, The driver model of the following vehicle is: Where a represents acceleration, a max v is the maximum acceleration that the following vehicle can achieve without restrictions, and v is the current following speed of the following vehicle. des Under ideal conditions, without any other vehicles obstructing the view, the maximum speed that the driver of the following vehicle expects to reach is δ, where δ is the acceleration exponent, representing the driver's reaction intensity to the difference between the following speed and the expected speed; s * (v, Δv) represents the desired distance between the following vehicle and the vehicle in front, Δv represents the speed difference between the following vehicle and the vehicle in front, s0 refers to the minimum distance maintained between the following vehicle and the vehicle in front when stationary, T is the minimum time interval that the driver expects to maintain while following, and b is the deceleration speed that the driver feels comfortable with during deceleration. This is used to ensure that the desired spacing is not negative.
6. The lane-changing decision-making method according to claim 5, characterized in that, Before step S2, the method further includes: based on the driver model of the interactive vehicle, using fuzzy logic rules to identify the driver style of the interactive vehicle, so that the driver style of the vehicle is taken into account when calculating the expected benefits of the following vehicle.
7. The lane-changing decision-making method according to claim 6, characterized in that, When using fuzzy logic rules to perform driver style recognition on the interactive vehicle, the maximum acceleration parameter 'a' in the driver model of the interactive vehicle is used as the basis. max The initial driver style of the interactive vehicle is obtained by taking into account the following distance parameter T. Then, based on the initial driver style of the interactive vehicle, the observation duration observeT when the driver model was established is considered to obtain the final driver style of the interactive vehicle.
8. The lane-changing decision-making method according to claim 1 or 6, characterized in that, The expected benefits of the controlled vehicle include speed benefits, safety benefits, and following distance benefits, while the expected benefits of the vehicle behind in the target lane include speed benefits and following distance benefits.
9. The lane-changing decision-making method according to claim 1, characterized in that, When the controlled vehicle has lane-changing space on both its left and right sides, in step S2, two k-level mixed strategy games are played respectively between the controlled vehicle and the vehicle behind it in its left lane, and between the controlled vehicle and the vehicle behind it in its right lane. In step S3, the lane-changing decisions of the controlled vehicle corresponding to these two k-level mixed strategy games are obtained. Then, the total expected payoffs of the controlled vehicle and the vehicles behind it in the left and right lanes are compared under the two lane-changing decisions. The lane-changing decision corresponding to the lane with the highest total expected payoff is selected as the final lane-changing decision of the controlled vehicle.
10. The lane-changing decision-making method according to claim 1, characterized in that, In step S3, when the controlled vehicle's lane-changing decision is to change lanes, and the strategy of the vehicle behind in the target lane under this lane-changing strategy is not to yield, the total expected benefits of the lane-changing and yielding strategy combination and the non-lane-changing and non-yielding strategy combination are compared, and the lane-changing decision of the controlled vehicle is obtained according to the strategy combination with the higher total expected benefits.
Citation Information
Patent Citations
A lane-changing decision-making method based on game theory
CN113335282B