Lane change decision method, apparatus, and vehicle
By calculating acceleration gains and using game theory, combined with communication interaction or irrational models, the problem of inaccurate prediction of human driver behavior in lane-changing decisions by autonomous vehicles is solved, resulting in more reasonable lane-changing decisions and reduced collision risks.
Patent Information
- Application Number
- CN202110312738.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-24
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2041-03-24
AI Technical Summary
Autonomous vehicles struggle to accurately predict human driver behavior when making lane-changing decisions, leading to irrational lane-changing decisions and increasing the risk of collisions.
By calculating the acceleration gains of the target vehicle before and after lane changing, game theory and irrational models are used to predict the behavior of human drivers. Combined with communication interaction or simulated game, the game equilibrium solution is obtained to make reasonable lane changing decisions.
It improves the accuracy of lane change decisions, reduces the risk of collisions due to predictive uncertainty, and is suitable for scenarios where autonomous vehicles and human-driven vehicles coexist.
Smart Images

Figure CN115123227B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of automatic driving, and in particular to a lane changing decision method and device, and a vehicle having the device. BACKGROUND
[0002] During the automatic driving process of a vehicle, a lane changing decision needs to be made according to actual situations. At this time, environmental information can be perceived through sensors such as cameras, radars, and lidars, and a lane changing decision can be made according to the perceived environmental information. The environmental information as a perception result is already occurred, and the decision needs to consider the states of the ego vehicle and a target vehicle in the future. Therefore, the ego vehicle needs to predict the future behavior of the target vehicle.
[0003] However, the prediction of the future behavior of the target vehicle by the ego vehicle is uncertain. In particular, in addition to the automatic driving vehicle on the road, there are also human-driven vehicles, and it is more difficult to predict the behavior of human drivers than to predict the behavior of automatic driving vehicles, because the behavior of humans is not always rational, and different people have different sensitivities to risks and benefits, resulting in different people making completely different decisions in the same scenario. This uncertainty may lead to an unreasonable lane changing decision by the ego vehicle, increasing the risk of collision. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a technology that can make a more reasonable lane changing decision and reduce the risk of collision.
[0005] To achieve the above purpose, the first aspect of the present application provides a lane changing decision method, comprising: obtaining a target lane into which a first vehicle intends to drive; calculating an acceleration benefit of a second vehicle in the case of lane changing of the first vehicle according to a first acceleration and a second acceleration of the second vehicle, the second vehicle being a vehicle on the target lane and located behind the first vehicle, the first acceleration being an acceleration of the second vehicle before the first vehicle implements a lane changing intention indicating behavior, and the second acceleration being an acceleration of the second vehicle after the first vehicle implements the lane changing intention indicating behavior; obtaining a game equilibrium solution according to the acceleration benefit of the second vehicle, the second vehicle being a vehicle on the target lane and located behind the first vehicle; and making a lane changing decision of the first vehicle according to the game equilibrium solution.
[0006] According to the lane changing decision method, the acceleration gain of the second vehicle is calculated according to the acceleration of the second vehicle before and after the first vehicle performs the lane changing intention indicating behavior, the calculated gain is closer to the real gain of the second vehicle, and thus the behavior of the second vehicle can be more accurately predicted, and the lane changing decision of the first vehicle can be more reasonable, and the collision risk caused by prediction uncertainty can be reduced.
[0007] As a possible implementation manner of the first aspect, an aggression value is calculated according to the first acceleration and the second acceleration, the aggression value indicates the aggression of the second vehicle, and the greater the aggression value, the smaller the acceleration gain of the second vehicle.
[0008] In this way, the gain of the second vehicle with a greater aggression value is set to be smaller, which can accurately reflect that the second vehicle is more sensitive to deceleration and less likely to yield, and thus the behavior of the second vehicle can be more accurately predicted, the lane changing decision of the first vehicle can be more reasonable, and the collision risk caused by prediction uncertainty can be reduced.
[0009] As a possible implementation manner of the first aspect, when the first acceleration and the second acceleration are both greater than a preset value, the aggression value is greater than or equal to the aggression value when one of the first acceleration and the second acceleration is less than the preset value.
[0010] In this way, when the first acceleration and the second acceleration of the second vehicle are both greater than a preset value, the second vehicle has no intention to decelerate even if the first vehicle shows a lane changing intention, and by setting a greater aggression value for the second vehicle, the gain of the second vehicle is smaller, the behavior of the second vehicle can be more accurately predicted, the lane changing decision of the first vehicle can be more reasonable, and the collision risk caused by prediction uncertainty can be reduced.
[0011] As a possible implementation manner of the first aspect, the acceleration gain of the second vehicle is calculated according to the first acceleration, the second acceleration, and a collision probability of the first vehicle and the second vehicle, and the smaller the collision probability, the greater the absolute value of the acceleration gain of the second vehicle.
[0012] In this way, the gain of the second vehicle is calculated according to the collision probability, and thus the gain of the second vehicle can be calculated according to the sensitivity of the second vehicle to risk, the behavior of the second vehicle can be accurately predicted, and the collision risk caused by prediction uncertainty can be reduced.
[0013] As a possible implementation manner of the first aspect of the present application, when the second vehicle does not support communication interaction, the acceleration benefit of the second vehicle under the lane-changing condition of the first vehicle is calculated according to the first acceleration and the second acceleration, and the game equilibrium solution is obtained according to the acceleration benefit of the second vehicle.
[0014] In this way, even if the second vehicle does not support communication interaction, the equilibrium solution can be obtained by simulating the game, and a reasonable lane-changing decision can be made, so as to cope with the scenario where the automatic driving vehicle and the human driving vehicle exist on the road at the same time.
[0015] As a possible implementation manner of the first aspect of the present application, the acceleration benefit of the second vehicle is calculated using the following benefit function U(a):
[0016]
[0017] wherein a is the acceleration of the second vehicle, w(a) is a weight function, is a value function, and w(a) is as follows:
[0018]
[0019] wherein η is a constant, P crash (a) is the collision probability of the second vehicle and the first vehicle corresponding to the acceleration a, P crash (a) is as follows:
[0020]
[0021] wherein τ is a prediction time length, Δd is the longitudinal distance between the first vehicle and the second vehicle, is an estimation of the speed of the first vehicle, which is subject to a Gaussian distribution with a mean of 0 and a variance of σ, and v is the current speed of the second vehicle,
[0022] as follows:
[0023]
[0024] wherein ∈ is a constant, λ is the aggression value, and λ is as follows:
[0025]
[0026] Where t is the time before the first vehicle performs the lane change intention indication behavior, a(t) is the first acceleration, t+1 is the time after the first vehicle performs the lane change intention indication behavior, a(t+1) is the second acceleration, and α is a constant greater than 0. For example, α is 2, and the range of λ is, for example, λ∈(0,2).
[0027] As a possible implementation of the first aspect of this application, when the second vehicle supports communication interaction, the acceleration gain of the second vehicle is calculated based on the information of the intention sequence of the second vehicle at n future time moments, the decision sequence information of the first vehicle at n future time moments is selected and sent to the second vehicle, the acceleration gain of the second vehicle is updated based on the information of the intention sequence sent by the second vehicle multiple times, and the game equilibrium solution is obtained based on the updated acceleration gain.
[0028] This method, which uses communication interaction to exchange intentions with the second vehicle to complete lane-changing decisions, can effectively reduce collision risk and prediction uncertainty. Furthermore, based on the idea of rolling time-domain optimization, the decision sequence for the next n time moments is optimized and sent each time, which can reduce decision oscillations and accelerate algorithm convergence.
[0029] As a possible implementation of the first aspect of this application, the decision sequence includes strategies E1 and E2 of the first vehicle, and the intention sequence includes strategies F1 and F2 of the second vehicle. The number of occurrences of the strategy combinations {E1, F1}, {E1, F2}, {E2, F1}, and {E2, F2} of the first vehicle and the second vehicle are calculated, and the acceleration gain of the second vehicle is calculated and updated based on the number of occurrences.
[0030] Using this method, only the target vehicle's decision-making information is needed during information exchange, without needing to know the other party's detailed benefit information, thus reducing the amount of information transmitted.
[0031] As one possible implementation of the first aspect of this application, the lane change intention indication behavior is to activate the turn signal or to laterally deviate in the direction of the target lane within the driving lane.
[0032] To achieve the above objectives, a second aspect of this application provides a lane-change decision-making device, comprising an acquisition module and a decision-making module. The acquisition module is used to acquire the target lane that a first vehicle intends to enter. The decision-making module is used to: calculate the acceleration gain of the second vehicle in the case of the first vehicle changing lanes, based on a first acceleration and a second acceleration of the second vehicle, wherein the second vehicle is a vehicle in the target lane and is located behind the first vehicle; the first acceleration is the acceleration of the second vehicle before the first vehicle expresses its intention to change lanes, and the second acceleration is the acceleration of the second vehicle after the first vehicle expresses its intention to change lanes; obtain a game equilibrium solution based on the acceleration gain of the second vehicle; and make a lane-change decision for the first vehicle based on the game equilibrium solution.
[0033] By employing the lane change decision-making device provided in the second aspect of this application, the acceleration gain of the second vehicle is calculated based on the acceleration of the second vehicle before and after the first vehicle performs the lane change intention expression behavior. The calculated gain is closer to the actual gain of the second vehicle, thereby enabling more accurate prediction of the behavior of the second vehicle, making more reasonable lane change decisions for the first vehicle, and reducing the collision risk caused by prediction uncertainty.
[0034] As a possible implementation of the second aspect of this application, the decision module calculates an aggression value based on the first acceleration and the second acceleration, the aggression value indicating the aggression of the second vehicle, and the larger the aggression value, the smaller the benefit of the second vehicle.
[0035] By setting the benefit of the second vehicle with a higher aggression value to a lower value, this approach can accurately reflect that the second vehicle is more sensitive to deceleration and less likely to give way. This allows for more accurate prediction of the second vehicle's behavior, more reasonable lane-changing decisions for the first vehicle, and a reduction in the risk of collision due to prediction uncertainty.
[0036] As one possible implementation of the second aspect of this application, the aggression value when both the first acceleration and the second acceleration are above a preset value is greater than or equal to the aggression value when either the first acceleration or the second acceleration is less than a preset value.
[0037] Using this method, when the first and second accelerations of the second vehicle are both above the preset values, even if the first vehicle shows an intention to change lanes, the second vehicle does not show any intention to slow down. By setting a larger aggression value for the second vehicle and making its benefits smaller, the behavior of the second vehicle can be predicted more accurately, and the lane-changing decision of the first vehicle can be made more reasonably, reducing the collision risk caused by prediction uncertainty.
[0038] As a possible implementation of the second aspect of this application, the decision module calculates the acceleration gain of the second vehicle based on the first acceleration, the second acceleration, and the collision probability between the first vehicle and the second vehicle. The smaller the collision probability, the greater the absolute value of the acceleration gain of the second vehicle.
[0039] This method also calculates the benefits of the second vehicle based on the collision probability, thus taking into account the second vehicle's sensitivity to risk when calculating its benefits, accurately predicting the behavior of the second vehicle, and reducing the collision risk caused by prediction uncertainty.
[0040] As a possible implementation of the second aspect of this application, when the second vehicle does not support communication interaction, the decision module calculates the acceleration gain of the second vehicle in the case of the first vehicle changing lanes based on the first acceleration and the second acceleration, and obtains the game equilibrium solution based on the acceleration gain of the second vehicle.
[0041] Using this method, even if the second vehicle does not support communication interaction, it can still find an equilibrium solution through simulated game theory and make a reasonable lane-changing decision, thus being able to cope with scenarios where autonomous vehicles and human-driven vehicles coexist on the road.
[0042] As a possible implementation of the second aspect of this application, when the second vehicle supports communication interaction, the decision module calculates the acceleration gain of the second vehicle based on the information of the intention sequence of the second vehicle at n future times, selects the decision sequence information of the first vehicle at n future times, and sends it to the second vehicle. The decision module updates the acceleration gain of the second vehicle based on the information of the intention sequence sent by the second vehicle multiple times, and obtains the game equilibrium solution based on the updated acceleration gain.
[0043] This method, which uses communication interaction to exchange intentions with the second vehicle to complete lane-changing decisions, can effectively reduce collision risk and prediction uncertainty. Furthermore, based on the idea of rolling time-domain optimization, the decision sequence for the next n time moments is optimized and sent each time, which can reduce decision oscillations and accelerate algorithm convergence.
[0044] As a possible implementation of the second aspect of this application, the decision sequence includes strategies E1 and E2 of the first vehicle, and the intention sequence includes strategies F1 and F2 of the second vehicle. The decision module calculates the number of occurrences of each strategy combination {E1, F1}, {E1, F2}, {E2, F1}, and {E2, F2} of the first vehicle and the second vehicle, and calculates and updates the revenue of the second vehicle based on the number of occurrences.
[0045] Using this method, only the target vehicle's decision-making information is needed during information exchange, without needing to know the other party's detailed benefit information, thus reducing the amount of information transmitted.
[0046] As one possible implementation of the second aspect of this application, the lane change intention indication behavior is to activate the turn signal or to laterally deviate in the direction of the target lane within the driving lane.
[0047] To achieve the above objectives, a third aspect of this application provides a vehicle having the lane change decision device described in any of the second aspect and its possible implementations.
[0048] To achieve the above objectives, a fourth aspect of this application provides a computing device, comprising: at least one processor; and at least one memory storing program instructions, wherein the at least one processor executes the program instructions to perform any of the methods described in the first aspect and its possible implementations.
[0049] To achieve the above objectives, a fifth aspect of this application provides a computer-readable storage medium having program instructions stored thereon, which a computer executes to perform any of the methods described in the first aspect and its possible implementations. Attached Figure Description
[0050] Figure 1 This is a schematic diagram illustrating an application scenario of an embodiment of this application;
[0051] Figure 2 This is a schematic diagram of the structure of a vehicle provided in an embodiment of this application;
[0052] Figure 3 This is a schematic diagram of a scenario where a target lane is selected in an embodiment of this application;
[0053] Figure 4 This is a schematic diagram of a scenario where a game vehicle is selected in an embodiment of this application;
[0054] Figure 5 This is a schematic diagram of the vehicle revenue function provided in an embodiment of this application;
[0055] Figure 6 This is a schematic diagram of the benefit function for the following vehicle in the adjacent lane provided in an embodiment of this application;
[0056] Figure 7 This is a schematic diagram of the value function curve provided in the embodiments of this application;
[0057] Figure 8 This is a schematic diagram of the weighting function curve provided in the embodiments of this application;
[0058] Figure 9This is a schematic diagram illustrating a scenario for calculating the aggression of a vehicle behind in an adjacent lane, as described in this application embodiment.
[0059] Figure 10 This is a schematic diagram of a scenario in which the lane-changing behavior of a vehicle following in an adjacent lane is predicted in an embodiment of this application.
[0060] Figure 11 This is a schematic flowchart of the automatic lane change decision-making method provided in the embodiments of this application;
[0061] Figure 12 This is a flowchart illustrating the information interaction game processing in the embodiments of this application;
[0062] Figure 13 This is a schematic diagram of the process for handling irrational model prediction intent in the embodiments of this application;
[0063] Figure 14 This is a schematic diagram of a lane-changing scenario used to compare the effects of lane-changing decision-making methods in the embodiments of this application with those in the prior art;
[0064] Figure 15 Is Figure 14 A diagram illustrating the use of rule-based methods in existing technologies to make lane-changing decisions in a given scenario;
[0065] Figure 16 Is Figure 14 A schematic diagram illustrating the use of the method described in this application to make a lane-changing decision in a given scenario;
[0066] Figure 17 This is a diagram used to illustrate the lane change decision-making method of existing technology;
[0067] Figure 18 This diagram illustrates the lane change decision-making methods used in existing technologies. Detailed Implementation
[0068] The technology provided in this application is used for autonomous vehicles to make automatic lane-changing decisions. Before specifically describing the technology provided in this application, please refer to... Figure 17 and Figure 18 A simple analysis of the existing technology is conducted.
[0069] In existing technologies, regarding automatic lane change decision-making technology, there is a rule-based method, referencing... Figure 17 This method is described below. When deciding whether to change lanes, vehicle Ego senses its own current acceleration 'a' and that of related vehicles (including vehicle A behind vehicle Ego in its target lane and vehicle B behind vehicle Ego in its current lane). ego ,a A ,a B Predict the acceleration of itself and related vehicles after changing lanes. A lane change is determined when both of the following conditions are met:
[0070] (1) Safety rule: After vehicle Ego changes lanes, the deceleration of the following vehicle A in the target lane must not exceed a given safety threshold b. safe ,Right now
[0071] (2) Expected condition: After vehicle Ego changes lanes, the sum of the acceleration gains of all vehicles is greater than a certain threshold Δa. th ,Right now
[0072] As mentioned above, in rule-based methods, the vehicle Ego determines whether to change lanes by sensing the current acceleration of the vehicle behind it, predicting its own and related vehicles' expected acceleration after the lane change, and setting safety and expectation rules. This method has the following drawbacks: First, it relies on the perceived current acceleration of related vehicles and the predicted acceleration after the lane change, resulting in significant uncertainty and requiring extremely high accuracy in perception and prediction; second, it does not consider irrational driver behavior, and the predicted acceleration fails to reflect actual driving behavior.
[0073] In existing technologies, there is also a method that uses game theory, as referenced Figure 18 The method is explained. For example... Figure 18 As shown, this method first selects a player vehicle. When changing lanes to an adjacent lane, the following vehicle in that lane is chosen as the player vehicle. A two-player game model is established between the player vehicle and the player vehicle, requiring both vehicles to support V2V (Vehicle-to-Vehicle) communication. The payoff table of the other vehicle is obtained through V2V communication. Then, a game theory algorithm is used to search for a Nash equilibrium. Once both vehicles reach a consensus, the strategy is executed. The payoff table is designed based on collision risk. This method has the following drawbacks: First, it requires both vehicles to have communication modules, making it only suitable for interactive decision-making between autonomous vehicles; second, the amount of information exchanged each time the payoff table is sent via V2V communication is large, and the computational cost of finding the optimal Nash equilibrium is significant; third, the decision results between adjacent decision cycles exhibit oscillations, and the lack of a backoff mechanism for failed lane changes poses a significant risk.
[0074] Due to the shortcomings of the existing technology, the automatic lane-changing decisions made may be unreasonable, which may increase the risk of collision.
[0075] In view of the above-mentioned problems of the prior art, the present application provides a technology that can make more reasonable automatic lane change decisions and reduce the risk of collision.
[0076] The following reference Figures 1 to 16 The embodiments of this application will be described in detail.
[0077] Figure 1 This is a schematic diagram illustrating an application scenario of an embodiment of this application. For example... Figure 1 As shown, vehicle 100 is driving on a road in autonomous driving mode. This road could be, for example, a city road or a highway. Imagine a scenario where vehicle 100 wants to change lanes to an adjacent lane and travel between other vehicles 200a and 200b.
[0078] When other vehicles 200a and 200b are vehicles that support communication interaction, such as autonomous vehicles, vehicle 100 can send its own intent information and receive the intent information of other vehicles 200a and 200b through its communication devices, and make decisions based on the intents of other vehicles 200a and 200b. When other vehicles 200a and 200b are not vehicles that do not support communication interaction, such as human-driven vehicles, vehicle 100 makes decisions based on environmental information and traffic information perceived by its sensing devices.
[0079] Figure 2 This is a structural diagram of vehicle 100. (For example...) Figure 2 As shown, vehicle 100 has a vehicle control device 30. Vehicle 100 also has a sensing system 10, a communication system 20, a power system 40, a steering system 50, and a braking system 60. In addition, vehicle 100 has structural elements other than these structural elements, but these are omitted here.
[0080] The perception system 10 includes sensors such as radar, cameras, and lidar. It also includes GPS (Global Positioning System), high-precision maps, and INS (Inertial Navigation System). The perception system 10 can acquire information about other vehicles and lane markings around the vehicle 100 using sensors such as radar, cameras, and lidar. Other vehicles mainly include vehicles in adjacent lanes and adjacent lanes. The perception system 10 can also acquire the vehicle 100's location information via GPS and high-precision maps, and its attitude information via INS.
[0081] The communication system 20 is capable of wireless communication with external objects (not shown). These external objects may include, for example, base stations (not shown), cloud servers, mobile terminals (smartphones, etc.), roadside equipment, other vehicles, etc. The communication system 20 is capable of V2V communication with other vehicles. For example, the communication system 20 can exchange rapidly changing dynamic information (such as location, speed, direction of travel, traffic conditions, etc.) with other vehicles via V2V communication using a PC5 interface. Additionally, as a vehicle-to-vehicle communication interface, the communication system 20 may also have a DSRC (Dedicated Short Range Communication) interface.
[0082] The powertrain 40 includes a drive ECU (Electronic Control Unit) (not shown) and a drive source (not shown). The drive ECU controls the driving force (torque) of the vehicle 100 by controlling the drive source. Examples of drive sources include an engine, a drive motor, etc. The drive ECU can control the drive source based on the driver's operation of the accelerator pedal, thereby controlling the driving force. Additionally, the drive ECU can also control the drive source based on commands sent from the vehicle control unit 30, thereby controlling the driving force. The driving force of the drive source is transmitted to wheels (not shown) via a transmission (not shown), thereby propelling the vehicle 100.
[0083] The steering system 50 includes a steering ECU (not shown) namely an EPS (Electric Power Steering) ECU and an EPS motor (not shown). The steering ECU can control the EPS motor according to the driver's operation of the steering wheel, thereby controlling the direction of the wheels (specifically the steering wheels). Additionally, the steering ECU can also control the EPS motor according to commands sent from the vehicle control unit 30, thereby controlling the direction of the wheels. Furthermore, steering can also be performed by changing the torque distribution or braking force distribution to the left and right wheels.
[0084] The braking system 60 includes a braking ECU (not shown) and a braking mechanism (not shown). The braking mechanism operates the braking components via a brake motor, hydraulic mechanism, etc. The braking ECU can control the braking mechanism based on the driver's operation of the brake pedal, thereby controlling the braking force. Additionally, the braking ECU can also control the braking mechanism based on commands sent from the vehicle control unit 30, thereby controlling the braking force. If the vehicle 100 is an electric vehicle or a hybrid vehicle, the braking system 60 may also include an energy recovery braking mechanism.
[0085] The vehicle control device 30 can be implemented by a single ECU or by a combination of multiple ECUs. An ECU is a computing device that includes a processor, memory, and communication interface connected via an internal bus. The memory stores program instructions, which, when executed by the processor, function as corresponding functional modules and units. The vehicle control device 30 is an example of the lane change decision-making device in this application.
[0086] The vehicle control device 30 has functional modules including an information acquisition module 31, a target lane selection module 32, a game-theoretic vehicle selection module 33, an information interaction game module 34, an irrational model prediction intent module 35, and a decision output module 36. The information acquisition module 31 and the target lane selection module 32 are examples of acquisition modules in this application. The information interaction game module 34 and the irrational model prediction intent module 35 are examples of decision modules in this application. The vehicle control device 30 implements these functional modules and / or functional units by executing programs (software) through a processor. However, the vehicle control device 30 can also implement all or part of these functional modules and / or functional units through hardware such as LSI (Large Scale Integration) and ASIC (Application Specific Integrated Circuit), or it can also implement all or part of these functional modules and / or functional units through a combination of software and hardware.
[0087] The information acquisition module 31 is responsible for acquiring input information from the perception system 10. The input information includes: the state information of the vehicle 100, including speed and acceleration; the relative motion state information of other vehicles around the vehicle 100, including relative position, speed, and acceleration; map information and vehicle positioning information. The information acquisition module 31 can preprocess the acquired input information. Preprocessing includes data interpolation, smoothing filtering, and removal of abnormal signals. The relative motion state information of other vehicles can be obtained not only through perception by the perception system 10, but also based on information received by the communication system 20 from the target vehicle, such as the target vehicle's position and speed, or by fusing the information perceived by the perception system 10 and the information received by the communication system 20. The information acquisition module 31 outputs the preprocessed input information to the target lane selection module 32.
[0088] The target lane selection module 32 extracts other vehicles within a certain range around vehicle 100 as target vehicles, and obtains relevant information about the target vehicles, including lane, position, speed, acceleration, etc. The target lane selection module 32 selects the target lane for automatic lane changing based on the following three indicators:
[0089] (1) The degree to which the vehicle's expected speed matches the average speed of the target lane;
[0090] (2) The vehicle spacing in the target lane and the inter-vehicle distance need to be large enough, that is, the risk is small enough;
[0091] (3) Mandatory lane change requirements, such as lanes that cannot be passed ahead in the current lane or lane change requirements for the planned route.
[0092] The target lane selection module 32 sets costs for each of the three indicators mentioned above, and selects the target lane based on the total cost. Below, refer to... Figure 3 The three-lane scenario illustrates the costs in these three aspects.
[0093] First, the speed cost g1. The speed cost g1 corresponds to the degree to which the vehicle's desired speed matches the average speed of the target lane. The vehicle's desired speed and the average speed of lane L... i The more mismatched the current average speed, the less likely that lane should be selected as the target lane, thus incurring a higher cost. Lane L i The speed cost g1 is expressed by the following equation (1):
[0094]
[0095] Where, d max It is a constant; Lane L i The current average speed of the target vehicle on the road; v ego d represents the vehicle's speed. max For example, it could be 100m.
[0096] exist Figure 3 In the scenario where vehicle 100 is traveling in lane L0, there are other vehicles 200c and 200d in the adjacent lane L1, which are the target vehicle. At this time, the speed information of other vehicles 200c and 200d can be used to calculate the average speed of other vehicles 200c and 200d, and then substituted into equation (1) to obtain the speed cost g1 of lane L1.
[0097] Second, the risk cost g2. Risk cost g2 corresponds to the vehicle spacing in the target lane. Vehicle spacing can be characterized by time headway (THW). Time headway refers to the time interval between the front ends of two consecutive vehicles passing a cross section in a convoy traveling in the same lane. To ensure safety, the time headway needs to be sufficiently large. The time headway reflects the level of risk; vehicles tend to choose lanes with lower risk. Lane L i The risk cost g2 is expressed by the following equation (2):
[0098]
[0099] in, For lane L i The number of target vehicles, THW i Lane L i The headway of the train.
[0100] Third, the necessity cost g3. The necessity cost g3 corresponds to whether there is a mandatory lane change requirement, such as when the current lane is impassable or the navigation route requires a lane change. The necessity cost g3 is expressed by the following equation (3):
[0101]
[0102] in, Lane L i The remaining distance; d max It is a constant.
[0103] For example, in Figure 3 In this scenario, vehicle 100 is traveling in lane L0, which is about to merge with the adjacent lane L1, thus requiring a mandatory lane change. At this point, the remaining distance in lane L0 can be used as a guide. To calculate the necessary cost g3. Remaining distance. It refers to the distance from the location of vehicle 100 to the point where lane L0 and the adjacent lane L1 meet.
[0104] Considering the costs from the above three aspects, including the vehicle's expectations, safety, and necessity, for any lane L... i Total Cost (L) i The following equation (4) represents:
[0105] Cost(L i )=ω1g1+ω2g2+ω3g3 (4)
[0106] Where ω1>0, ω2<0, and ω3>0 are all weights.
[0107] According to equation (4) above, the total cost Cost(L) is selected. i The smallest lane is selected as the target lane.
[0108] The game vehicle selection module 33 selects target vehicles that may affect vehicle 100's lane change as the game object, i.e., the game vehicle. Vehicle 100 is an example of the first vehicle in this application. The game vehicle is an example of the second vehicle in this application. Typically, the game vehicle selection module 33 selects target vehicles in adjacent lanes and adjacent adjacent lanes within a specified longitudinal distance as game vehicles. Here, the adjacent adjacent lane of the current lane refers to the adjacent lane of the adjacent lane of the current lane. For example, in Figure 4In this scenario, the current lane is lane L2, the adjacent lanes are lanes L1 and L3, and the lanes immediately following the current lane are lanes L0 and L4. The specified distance can be set based on vehicle speed. For example, the specified distance can be set to 100m. This considers target vehicles within 100m behind the current vehicle in adjacent lanes and the lanes immediately following the current vehicle.
[0109] exist Figure 4 In this scenario, when vehicle 100, traveling in lane L2, intends to change lanes to the left (i.e., choosing adjacent lane L3 as the target lane), target vehicle 200b in adjacent lane L3 and target vehicle 200a in adjacent lane L4 are chosen as the playing vehicles. This is because the lane-changing behavior of target vehicle 200a and the acceleration / deceleration behavior of target vehicle 200b will affect vehicle 100's lane-changing behavior. Here, we can assume that target vehicles 200e and 200f ahead of vehicle 100 are driving normally in their current state and their influence is temporarily ignored. Furthermore, since adjacent lane L3 is chosen as the target lane, target vehicles 200c and 200d in lanes L0 and L1 will not affect vehicle 100's lane-changing behavior.
[0110] When the game vehicle selection module 33 selects a game vehicle that supports communication interaction, the information interaction game module 34 exchanges intention information with the game vehicle through the communication system 20, finds the equilibrium solution based on game theory, obtains the optimal strategy, and makes an automatic lane change decision when a consensus is reached with the game vehicle.
[0111] The three elements of game theory include participants, strategies, and payoffs. Here, "participants" include vehicle 100 and the vehicles selected by the vehicle selection module 33. Vehicles that are participants in the game are divided into two types: vehicles following in adjacent lanes and vehicles following in adjacent lanes. For example, a vehicle following in an adjacent lane... Figure 4 The target vehicle in the adjacent lane (L3) is vehicle 200b. The vehicle following in the adjacent lane is, for example... Figure 4 The target vehicle is 200a in lane L4 of the adjacent lane.
[0112] For vehicle 100 (the self-driving vehicle), the "strategy" in the above three elements includes immediate lane change and waiting to change lanes. Regarding the vehicle behind in the adjacent lane, if the vehicle behind in the adjacent lane chooses to decelerate, it means that vehicle 100 can safely change lanes; otherwise, there is a risk of collision. Therefore, vehicle 100 needs to consider the acceleration and deceleration strategies of the vehicle behind in the adjacent lane. Regarding the vehicle behind in the adjacent lane, since the vehicle behind in the adjacent lane is not in vehicle 100's target lane, its longitudinal acceleration and deceleration behavior will not directly affect vehicle 100's lane change. However, its lateral lane change behavior may bring a collision risk to vehicle 100's lane change. Therefore, vehicle 100 needs to consider the lateral lane change strategy of the vehicle behind in the adjacent lane. Thus, for the vehicle behind in the adjacent lane, the "strategy" in the above three elements includes acceleration and deceleration; for the vehicle behind in the adjacent lane, the "strategy" in the above three elements includes changing lanes and not changing lanes. It should be noted that the two strategies adopted by vehicle 100 are merely examples. Vehicle 100's strategies are not limited to two; there can be more, such as accelerating to change lanes, decelerating to change lanes, and waiting to change lanes.
[0113] The "payoff value" among the three factors mentioned above is a quantifiable measure of the effectiveness of different strategies. Here, we temporarily disregard the impact of vehicles in adjacent lanes and focus solely on the game between vehicle 100 and the vehicle in the adjacent lane. Vehicle 100 tends to merge into the target lane safely and quickly; therefore, merging time and collision risk are used as indicators to measure payoff. The vehicle in the adjacent lane, on the other hand, tends to maintain its original driving state safely; therefore, change in acceleration is used to measure payoff.
[0114] Reference Figure 5 The payoff function for vehicle 100 is explained. Using merging time and collision risk as metrics, the potential payoffs for vehicle 100 under different strategies are calculated. Figure 5 As shown, when vehicle 100 chooses to merge, there is a risk of collision if the vehicle behind in the adjacent lane accelerates or decelerates. Therefore, the benefit is the weighted sum of the merging time and the collision time: α·t merge +β·TTC, where t merge Let α be the merging time and TTC be the collision time (Time-To-Collision). α < 0 and β > 0 are both weighting coefficients. When vehicle 100 chooses to wait, there is no collision risk, so only the merging time is considered, and the payoff is: γ·t merge , where γ<0 is the weighting coefficient.
[0115] Reference Figure 6 The payoff function for the vehicle following in the adjacent lane is explained. Using the change in acceleration as a metric, the possible payoffs for the vehicle following in the adjacent lane when choosing different strategies are calculated. Figure 6As shown, when a vehicle in the adjacent lane accelerates, there is a risk of collision if the vehicle in the adjacent lane chooses to merge. To avoid this risk, the vehicle in the adjacent lane needs to brake suddenly, and the benefit is ω·Δa. brake Where ω is the weighting coefficient, Δa brake This represents the change in acceleration during emergency braking; if vehicle 100 chooses to wait, there is no risk of collision, and vehicles in the adjacent lane can proceed normally, with a benefit of ω·Δa. normal Where ω is the weighting coefficient, Δa normal This refers to the change in acceleration during normal driving. When a vehicle in the adjacent lane chooses to slow down and yield, regardless of whether the vehicle chooses to merge or wait, the benefit to the vehicle in the adjacent lane is defined as the change in acceleration, and the benefit is ω·Δa. yield Where ω is the weighting coefficient, Δa yield This refers to the change in acceleration during the expected deceleration. In summary, the gains of the vehicle following in the adjacent lane are all measured in terms of the change in acceleration. This describes the gain function used by the vehicle following in the adjacent lane to calculate its own gain. However, in this embodiment, vehicle 100 does not receive such a calculated gain table from the vehicle following in the adjacent lane, which is acting as the game vehicle. Instead, it only receives the intention information of the game vehicle and estimates the gain of the game vehicle based on this intention information.
[0116] When the game vehicles support communication interaction, the information interaction game module 34 uses... Figure 5 The payoff function of vehicle 100 is used to calculate its own payoff, determine the decision sequence of vehicle 100 at n future times, and send it to the playing vehicle via communication system 20. Then, vehicle 100 receives the intention sequence of the playing vehicle at n future times from the playing vehicle. Based on the intention sequence made by the playing vehicle after receiving vehicle 100's decision sequence, the payoff of the playing vehicle is estimated using statistical methods to obtain a payoff table. Based on the payoff table, vehicle 100's payoff is calculated, and the decision sequence at the n future times with the highest payoff is selected and sent to the playing vehicle. The information interaction game module 34 repeatedly performs information interaction and payoff estimation with the playing vehicle. When a consensus is reached with the playing vehicle, a game equilibrium solution is obtained. Then, the information interaction game module 34 makes an automatic lane-changing decision based on the game equilibrium solution.
[0117] When the vehicle selected by the vehicle selection module 33 does not support communication interaction, for example, when the vehicle is a human-driven vehicle, the irrational model intention prediction module 35 uses an irrational model to estimate the payoff of the adjacent vehicle in the next lane under different strategies. Based on this payoff, it predicts the intention of the adjacent vehicle, simulates the game process in game theory, searches for an equilibrium solution, and makes an automatic lane-changing decision. In this case, similarly to the above, the payoff of vehicle 100 uses... Figure 5 The profit function shown is used for calculation. In this embodiment, the value function of the designed irrational model is... The decision weights w(a) and the payoff function U(a) are as follows:
[0118]
[0119] Where λ is the aggression factor; a is the acceleration; and ∈ is a constant. ∈ is an empirical parameter whose value is predetermined based on measured data. The range of values for ∈ is, for example, ∈ ∈ (0, 1).
[0120]
[0121] Where η is a constant; P crash (a) represents the collision probability corresponding to acceleration *a*. η is an empirical parameter whose value is predetermined based on measured data. The range of η is, for example, η∈(0,1). According to the CA (Constant acceleration) model, P... crash (a) can be obtained from the following formula:
[0122]
[0123] Where τ is the prediction time length, Δd is the longitudinal distance between the two workshops, and a is the acceleration; The speed of the vehicle ahead is estimated by a Gaussian distribution with a mean of 0 and a variance of σ; v is the current speed. τ and σ are pre-calibrated parameters.
[0124]
[0125] The aggression factor λ in equation (5) above represents the driver's aggression. Vehicle 100 can demonstrate its lane-changing intention to the other vehicle by performing actions that indicate its intention to change lanes. Based on the longitudinal acceleration of the other vehicle before and after such actions, the aggression factor of the other vehicle is calculated by the irrational model intention prediction module 35. Actions that indicate lane-changing intentions include, for example, vehicle 100 activating its turn signal or shifting laterally in the direction of the target lane within its driving lane for a certain period of time. At this time, the perception system 10 perceives the acceleration a of the i-th (i = 1, 2, ..., m, m ≥ 1) vehicle behind in the adjacent lane at time t before vehicle 100 performs the aforementioned actions and at time t+1 after performing the aforementioned actions. i (t), a i (t+1), the aggressiveness factor λ of the game vehicle is calculated by the irrational model prediction intention module 35 according to the following formula. i . λ i The range of values for is, for example, λ. i∈(0,2). Regarding the setting of time t and t+1, for example, time t can be set as the time when vehicle 100 is about to perform the above behavior, and time t+1 can be the time after a certain period of time has elapsed since the start of the above behavior. The certain period of time can be, for example, 2 seconds, or other values.
[0126]
[0127] Equation (9) is based on the payoff function and the observed changes in the acceleration of the vehicles in the game, calculating and estimating the aggression factor that best matches the model. According to Equation (9), when a i (t)<0 or a i When (t+1)<0, use a i (t) and a i (t+1) Calculate the aggression factor λ; otherwise, the game vehicle does not decelerate, is considered to have high aggression, and the aggression factor λ is directly set to 2. In the embodiments of this application, a i (t) and a i The aggression factor λ when both sides are greater than 0 (t+1) is greater than or equal to the aggression factor λ when one side is less than 0. The aggression factor λ is an example of the "aggression value" in this application. i (t) is an example of the "first acceleration" in this application, a i (t+1) is an example of the "second acceleration" in this application, and 0 is an example of the "preset value" in this application. Furthermore, λ i The value "2" in the range of values is just one example; its maximum value can also be any other value greater than 0.
[0128] Driver aggression has a significant impact on driving behavior. A large aggression factor λ indicates that the driver is more sensitive to deceleration and less likely to yield; a small aggression factor λ indicates that the driver is more sensitive to risk and more likely to yield. The calculated aggression factor λ is fitted to the value function. Value function The magnitude depends on the acceleration a and the aggression factor λ. Figure 7 This represents the value function curves corresponding to different aggression factors λ. For example... Figure 7 As shown, when a < 0, the larger the aggression factor, the steeper the curve, and the greater the change in the value function with acceleration a, corresponding to the driver's greater sensitivity to changes in speed. When a < 0, under the same acceleration, the larger the aggression factor λ, the greater the value function... The smaller the value of λ, the smaller the acceleration, given the same aggression factor λ. The smaller the value of a, the more independent it is of the aggression factor λ when a ≥ 0. The smaller the acceleration a, the more the value of the function... The smaller the value.
[0129] From equation (8), it can be seen that the value function The magnitude of the aggression factor λ is related to the estimated payoff of the game vehicle. In this embodiment, the larger the aggression factor λ is, the smaller the payoff of the game vehicle calculated according to equation (8).
[0130] Furthermore, according to equation (8), the payoff of the game vehicle estimated by the irrational model prediction intention module 35 is also related to the magnitude of the decision weight w(a). As shown in equation (6), the decision weight w(a) is determined based on the collision probability P. crash (a) Calculated. Figure 8 Let w(a) be the function curve of the decision weights. Figure 8 In the diagram, the horizontal axis represents the collision probability P. crash (a), where the vertical axis represents the decision weight w(a). From Figure 8 It can be known that the collision probability P crash (a) The smaller the value of a, the larger the value of the decision weight w(a). Combining equations (5) and (8), it can be seen that when a ≥ 0, the payoff of the game vehicle is positive, and the larger the decision weight w(a), the larger the payoff of the game vehicle; when a < 0, the payoff of the game vehicle is negative, and the larger the decision weight w(a), the smaller the payoff of the game vehicle. That is, in the embodiment of this application, the collision probability P crash (a) The smaller the value, the greater the absolute value of the game vehicle's payoff. The payoff of the game vehicle calculated according to equation (8) is an example of the "acceleration payoff" in this application.
[0131] The above explains the irrational model used in the irrational model prediction intent module 35 to estimate the payoff of the vehicle in the adjacent lane as a game vehicle. For example, in Figure 9 In this scenario, vehicle 100 can first signal its intention to change lanes to target vehicle 200b by activating its turn signal or by shifting laterally within lane L0 towards target lane L1. It then senses the acceleration of target vehicle 200b before and after signaling its intention and substitutes this acceleration into equation (9) to calculate the aggression factor λ of the driver of target vehicle 200b. The calculated aggression factor λ is then substituted into the irrational model shown in equations (5) to (8). The irrational model is used to calculate the gains of target vehicle 200b under different accelerations and predict the intention of target vehicle 200b. Through simulated game theory, vehicle 100 makes an automatic lane-changing decision.
[0132] In addition, for Figure 9 The target vehicle 200a in the adjacent lane, i.e., lane L2, can be predicted to change lanes as follows. Figure 9In the example, target vehicles 200c, 200d, and 200f are vehicles with different speeds in the three lanes. It is assumed that target vehicle 200c has a slower speed and target vehicle 200d has a faster speed. If target vehicle 200a wants to avoid colliding with target vehicle 200c, there are only two options: (1) decelerate and stay in the current lane; (2) change lanes to lane L1, and the speed change does not need to be too large. The two options correspond to two strategies: deceleration and lane change. First, according to the CA model, the deceleration a1 required for target vehicle 200a to avoid colliding with target vehicle 200c is calculated; second, the deceleration a2 required for target vehicle 200a to avoid colliding with target vehicle 200d after changing lanes to lane L1 is calculated. If a2-a1 is greater than a certain threshold, it is considered that target vehicle 200a will change lanes. The threshold is, for example, an empirical value set in advance based on real data. In the case that target vehicle 200a, which is the following vehicle in the adjacent lane, will change lanes, it can be done as follows Figure 10 As shown, the benefit of vehicle 100 is calculated by assuming target vehicle 200a changes lanes to a corresponding position on lane L1. Figure 10 In the scenario shown, the probability of a collision increases, which reduces the lane-changing benefit of vehicle 100. Therefore, vehicle 100 will tend to choose to stay in the current lane.
[0133] The decision output module 36 generates control commands for sending to the power system 40, steering system 50 and braking system 60 based on the decision results of the information interaction game module 34 or the irrational model prediction intention module 35, so as to control the power system 40, steering system 50 and braking system 60 and make the vehicle 100 drive according to the decision made.
[0134] Next, refer to Figure 11 The specific steps of the automatic lane change decision method executed by the vehicle control device 30 are explained.
[0135] In step S10, the information acquisition module 31 acquires input information. The input information includes, for example, the state information of vehicle 100, information about surrounding target vehicles, lane line information, etc. The state information of vehicle 100 includes its position, speed, and acceleration. The target vehicle information mainly includes the relative position, speed, and acceleration of target vehicles in adjacent lanes and adjacent lanes. Optionally, the information acquisition module 31 can use interpolation methods to synchronize signals to obtain input information with the same update period, addressing the issue of different sensor sampling frequencies.
[0136] In step S20, the target lane selection module 31 selects the target lane for automatic lane changing based on the input information. Specifically, as described above, it considers three aspects of cost and selects the lane with the lowest total cost as the target lane. Note that this step only selects the desired lane and does not immediately perform the lane change.
[0137] After determining the target lane in step S20, the game vehicle selection module 33 selects the game vehicle in step S30. Specifically, as described above, a target vehicle that will affect the lane-changing process of vehicle 100 can be selected as the game vehicle. Here, the target vehicle located behind vehicle 100 in the target lane is selected as the game vehicle.
[0138] In step S40, it is determined whether the game vehicle selected in step S30 supports communication interaction. For example, the communication system 20 can send a communication request message to the game vehicle via V2V broadcast. This message may include information such as the vehicle 100's location and speed, and a request for a response from surrounding vehicles within a certain range. If a response message is received from the game vehicle within a specified time after sending the communication request message, the game vehicle is considered to support communication interaction; otherwise, it is considered not to support communication interaction. Alternatively, the location information of the game vehicle sensed by the sensing system 10 can be compared with the location information of other vehicles received via V2V communication. If the comparison result shows that the two location information are basically consistent, it is determined that a response message from the game vehicle has been received. If the game vehicle supports communication interaction, proceed to step S50; otherwise, proceed to step S60.
[0139] In step S50, the information interaction game module 34 performs the information interaction game processing. Details regarding the information interaction game processing will be explained later.
[0140] In step S60, the irrational model prediction intent module 35 performs irrational model prediction intent processing. Details regarding the irrational model prediction intent will be explained later.
[0141] In step S70, the decision result from step S50 or S60 is output. That is, based on the decision result from step S50 or S60, control commands are generated to be sent to the powertrain 50, steering system 60, and braking system 70, so that the vehicle 100 can drive according to the made decision. After step S70 ends, Figure 11 The series of processes shown has ended.
[0142] Figure 12 This is a schematic diagram illustrating the specific process of information interaction and game theory in step S50. (Refer to...) Figure 12 A detailed explanation of information interaction game processing is provided.
[0143] In step S51, at time t = t0, the decision sequence of vehicle 100 (the self-vehicle) for the next n time moments is sent to the playing vehicle. In this step, the information interaction game module 34 can predict the state of the playing vehicle for the next n time moments based on the position, speed, and acceleration information of the playing vehicle sensed by the sensing system 10 or obtained through V2V communication, and determine the decision sequence of vehicle 100 for the next n time moments. The decision sequence of vehicle 100 consists of the strategy corresponding to each time moment. The strategies corresponding to each time moment include, for example, waiting and joining. The value of n can be set according to the computing power of the vehicle control device 30. n can be set to 3, for example, but is not limited to this, as long as it is greater than 0. If the computing power allows, setting n to be larger can speed up the convergence of the algorithm.
[0144] In step S52, at time t = t + 1, information about the intention sequence sent by the game vehicle is received, and the game vehicle's payoff is estimated. The received intention sequence of the game vehicle contains its strategy for the next n time steps. The game vehicle's strategy includes, for example, acceleration and deceleration. In this embodiment, vehicle 100 does not obtain the game vehicle's payoff table from the game vehicle itself, but instead uses statistical methods to approximately estimate the game vehicle's payoff to obtain the game vehicle's payoff table.
[0145] Based on all the intention sequences (historical intention sequences) received so far and the decision sequences sent by vehicle 100 so far (historical decision sequences), the number of times each vehicle chooses various strategies when vehicle 100 selects a certain strategy is counted, thereby estimating the payoff of vehicle 100 under different choices. In other words, the payoff of vehicle 100 under different strategies is estimated by counting the number of times the strategies in vehicle 100's historical decision sequences and the corresponding strategies in the game vehicle's historical intention sequences occur. For example, if vehicle 100's strategies include waiting and merging, and the game vehicle's strategies include accelerating and decelerating, based on the corresponding strategies adopted by the game vehicle when vehicle 100 chooses any of these strategies, there are four possible strategy combinations between vehicle 100 and the game vehicle: {wait, accelerate}, {wait, decelerate}, {merge, accelerate}, and {merge, decelerate}. If vehicle 100's historical decision sequence includes a total of 10 instances of "waiting," and the strategy combination {wait, accelerate} appears 8 times—meaning the game vehicle chose "accelerate" 8 times out of the 10 instances of "waiting"—then the estimated payoff for game vehicle 100 choosing "accelerate" when vehicle 100 chooses "waiting" is 8 / 10. With... Figure 12As the number of iterations in the algorithm increases, both the numerator and denominator change, and the estimated payoff of the game vehicle eventually approaches the actual payoff. {Wait, Accelerate}, {Wait, Decelerate}, {Merge, Accelerate}, and {Merge, Decelerate} are examples of the strategy combinations {E1, F1}, {E1, F2}, {E2, F1}, and {E2, F2} for the first and second vehicles in this application, respectively. The payoff of the game vehicle estimated using the above statistical method is an example of the "acceleration payoff" in this application. Furthermore, the types of strategies for vehicle 100 and the game vehicle are not limited to the two mentioned above; they can also be three or more. In this case, the same statistical method can be used to calculate the payoff of the game vehicle.
[0146] In step S53, based on the estimated payoffs of the game vehicles, the payoff of vehicle 100 is calculated, and the decision sequence for the next n time steps that maximizes the payoff of vehicle 100 is selected and sent to the game vehicles. For example, the information interaction game module 34 can calculate the payoff of vehicle 100 for each strategy at each of the next n time steps, and the decision sequence that maximizes the sum of payoffs over the next n time steps is the best choice for vehicle 100.
[0147] In step S54, it is determined whether the two parties in the game have reached a consensus. If they have, the information exchange game processing ends; otherwise, it returns to step S51 and repeats steps S52-S54 until both parties reach a consensus. For example, a consensus can be considered reached when both parties' current strategies are consistent with their strategies at the previous moment. In this case, the decision sequence sent by vehicle 100 at the end is the decision result of vehicle 100.
[0148] Figure 13 This is a schematic diagram illustrating the specific process of processing the irrational model prediction intent in step S60. (Refer to...) Figure 12 A detailed explanation is provided regarding the handling of irrational model prediction intentions.
[0149] In step S61, the aggression factor of the playing vehicle is calculated. Vehicle 100 performs an action indicating its intention to change lanes to the playing vehicle, and interacts with the playing vehicle. Based on the longitudinal acceleration change of the playing vehicle, the aggression factor λ of the playing vehicle is calculated. The action indicating the intention to change lanes is, for example, vehicle 100 turning on its turn signal or shifting laterally in the direction of the target lane within the driving lane for a certain period of time. At this time, the perception system 10 perceives the acceleration a(t) and a(t+1) of the playing vehicle at time t before vehicle 100 performs the above action and at time t+1 after the above action, and calculates the aggression factor λ of the playing vehicle according to the above equation (9). In addition, the speed v of the playing vehicle and the longitudinal distance Δd between the playing vehicle and vehicle 100 in the equation (7) used at this time can be, for example, perceived by the perception system 10 after time t and before time t+1. At this time, vehicle 100 is equivalent to the vehicle in front referred to in equation (7). This is an estimate of the speed of vehicle 100.
[0150] In step S62, an irrational model is established. The aggression factor λ calculated in step S61 is substituted into equation (5) above to obtain the value function. According to this value function And the above equation (6) based on the collision probability P crash (a) Calculate the decision weight w(a), and obtain the payoff function U(a) of the game vehicle shown in Equation (8). Using the payoff function U(a), the payoff value of the game vehicle under different accelerations can be estimated.
[0151] In step S63, the irrational model established in step S62 is used to calculate the payoff of the playing vehicle and predict its intentions, thereby simulating the game and finding the equilibrium solution. Specifically, vehicle 100 can use the irrational model established in step S62 to estimate the payoff value of the playing vehicle under different conditions, and vehicle 100 according to... Figure 5The payoff function shown calculates the payoff of the vehicle under different conditions. Thus, vehicle 100 can obtain the payoff table of the two-player game. The strategy of vehicle 100 can be determined based on the payoff table of vehicle 100, such as waiting or merging. The strategy of the other vehicle can be predicted based on the payoff table of the other vehicle, such as accelerating or decelerating. Then, for example, the game process can be simulated by fictitious play in game theory algorithms to find the equilibrium solution. At the beginning of each round of the game, vehicle 100 and the other vehicle find an optimal strategy based on the other's historical strategy. Then, they update their own historical strategy based on the historical strategy and the optimal strategy of this round. As the iteration continues, the strategy will gradually converge to the Nash equilibrium. The equilibrium solution obtained in this way is the decision result of vehicle 100. In addition, when using the irrational model to calculate the payoff of the other vehicle, the speed v of the other vehicle and the longitudinal distance Δd between the other vehicle and vehicle 100 in the equation (7) are obtained in real time by the sensing system 10. It should be noted that the game theory algorithm used here is not limited to Fictitious Play, but can also be other existing game theory algorithms, such as the line drawing method.
[0152] Next, refer to Figures 14 to 16 Taking a scenario where the vehicles in the game do not support communication interaction as an example, the beneficial effects of the interactive automatic lane-changing decision-making provided in this application embodiment compared with the traditional rule-based method are explained.
[0153] Figure 14 This is a diagram illustrating a lane-changing scenario. (For example...) Figure 14 As shown in (a), vehicle 100 is traveling in lane L0 and has to change lanes to merge into adjacent lane L1. There are three target vehicles 200a, 200b and 200c behind vehicle 100 in adjacent lane L1. Figure 14 (b) represents the initial state of vehicle 100 and target vehicles 200a, 200b, and 200c, namely the speed of each vehicle and the longitudinal distance between target vehicles 200a, 200b, and 200c and vehicle 100.
[0154] Figure 15 This describes the scenario when changing lanes using traditional rule-based methods. For example... Figure 15 As shown, due to insufficient safe distance between target vehicles 200a, 200b, and 200c, vehicle 100 can only choose to change lanes to either in front of target vehicle 200a or behind target vehicle 200c. However, calculations show that if it wants to change lanes to in front of target vehicle 200a within 4 seconds, the required acceleration is 4 m / s². 2 Most vehicles cannot achieve this acceleration, so they can only slow down and wait for all target vehicles 200a, 200b, and 200c to pass before changing lanes. However, the edge of lane L0 is close ahead, and the situation is dangerous.
[0155] Figure 16 This describes a lane-changing scenario using the method described in this application embodiment. According to the method described in this application embodiment, vehicle 100 makes a decision based on an irrational model predicting the intentions of target vehicles 200a, 200b, and 200c. For example... Figure 16 As shown in (a), vehicle 100 activates its turn signal or laterally deviates within lane L0. The motion states of target vehicles 200a, 200b, and 200c are observed, and their aggression factors are calculated. The results show that target vehicle 200a exhibits greater aggression and is more sensitive to deceleration. Therefore, as shown in (a), vehicle 100... Figure 16 As shown in (b), vehicle 100 slows down to allow target vehicle 200a to pass. Furthermore, it was found that target vehicle 200b is less aggressive and more risk-sensitive. Using an irrational model to calculate its payoff under different strategies, it was concluded that when vehicle 100 changes lanes, target vehicle 200b will slow down to avoid it. Therefore, vehicle 100 merges in before target vehicle 200b, thus... Figure 16 As shown in (c), vehicle 100 merges into adjacent lane L1 before reaching the edge of its own lane L0, and travels between target vehicles 200a and 200b.
[0156] The automatic lane-changing decision-making method provided in this application, when the playing vehicle is an autonomous vehicle supporting V2V communication interaction, vehicle 100 and the playing vehicle exchange intention information via V2V communication. Interactive decision-making between vehicles is conducted using game theory methods, reaching a consensus (equilibrium solution) during the game process. This considers the mutual influence and interaction between vehicles, moving away from independent decision-making and reducing decision risk and uncertainty. Furthermore, vehicle 100 optimizes in the rolling time domain, optimizing and sending the decision sequence for the next n time moments each time, reducing decision oscillations and accelerating algorithm convergence. In addition, since only the intention information of the playing vehicle needs to be received during information interaction, and not detailed payoff information, the amount of information transmitted is reduced.
[0157] Using the automatic lane-changing decision-making method provided in this application embodiment, when the playing vehicle is a human-driven vehicle that does not support V2V communication, its intention information cannot be directly obtained. Therefore, vehicle 100 interacts with it by using turn signals, lateral deviation, etc., to estimate and quantify the aggressiveness of the playing vehicle, and then establishes an irrational model to predict its gains under different situations, simulates the game process, and finds the optimal solution. Thus, the established irrational model is more consistent with human driving behavior in real scenarios, which can reduce the prediction error of human driver behavior, thereby reducing decision-making risks and uncertainties.
[0158] The automatic lane-changing decision-making method provided in the embodiments of this application has been described above. However, the embodiments of this application are not limited to the above structure. In the above description, after selecting the game vehicle, it is determined whether the game vehicle supports communication interaction. If communication interaction is supported, information interaction game processing is performed; if communication interaction is not supported, irrational model prediction intention processing is performed. However, it is also possible to omit the determination of whether communication interaction is supported, and regardless of whether communication interaction is supported, irrational model prediction intention processing is performed to make an automatic lane-changing decision. In addition, the payoff function of vehicle 100 may not be used. Figure 5 The revenue function shown is different from other known revenue functions. In equations (5) and (6), ∈ and η can also take other ranges of values.
[0159] Based on the above description, the embodiments of this application are summarized as follows.
[0160] This application provides a lane change decision method, including: obtaining the target lane that a first vehicle (e.g., vehicle 100) intends to enter; and determining the lane change decision based on the second vehicle (e.g., vehicle 100). Figure 9 The game equilibrium is calculated by taking the first acceleration of the vehicle following the first vehicle (200b) in the adjacent lane before the first vehicle indicates its intention to change lanes, and the second acceleration of the second vehicle after the first vehicle indicates its intention to change lanes. The second vehicle is a vehicle in the target lane and is located behind the first vehicle. Based on the acceleration gain of the second vehicle, the game equilibrium solution is obtained. Based on the game equilibrium solution, the first vehicle makes a lane-changing decision. The lane-changing intention is indicated, for example, by activating a turn signal or by laterally shifting within the driving lane towards the target lane.
[0161] As shown above, the acceleration gain of the second vehicle is calculated based on the acceleration of the second vehicle before and after the first vehicle expresses its intention to change lanes. The calculated gain is closer to the actual gain of the second vehicle, which can more accurately predict the behavior of the second vehicle, make a more reasonable lane-changing decision for the first vehicle, and reduce the collision risk caused by prediction uncertainty.
[0162] Optionally, an aggression value is calculated based on the first acceleration and the second acceleration. The aggression value indicates the aggression of the second vehicle; the larger the aggression value, the smaller the acceleration benefit of the second vehicle. Therefore, setting a smaller benefit for a second vehicle with a larger aggression value accurately reflects that such a vehicle is more sensitive to deceleration and less likely to yield, thus enabling more accurate prediction of the second vehicle's behavior.
[0163] Optionally, the aggression value when both the first acceleration and the second acceleration are above a preset value is greater than or equal to the aggression value when either the first acceleration or the second acceleration is below a preset value. When both the first acceleration and the second acceleration of the second vehicle are above the preset value, even if the first vehicle shows an intention to change lanes, the second vehicle does not show any intention to decelerate. By setting a larger aggression value for the second vehicle, resulting in a smaller benefit, the behavior of the second vehicle can be predicted more accurately, and the lane-changing decision of the first vehicle can be made more rationally, reducing the collision risk caused by prediction uncertainty.
[0164] Optionally, the acceleration gain of the second vehicle is calculated based on the first acceleration, the second acceleration, and the collision probability between the first vehicle and the second vehicle. The lower the collision probability, the greater the absolute value of the acceleration gain of the second vehicle. This allows for the calculation of the gain taking into account the second vehicle's sensitivity to risk, accurately predicting the behavior of the second vehicle, and reducing the collision risk caused by prediction uncertainty.
[0165] Optionally, when the second vehicle does not support communication interaction, the acceleration gain of the second vehicle in the case of the first vehicle changing lanes is calculated based on the first acceleration and the second acceleration. The game equilibrium solution is then obtained based on the acceleration gain of the second vehicle. Therefore, even if the second vehicle does not support communication interaction, an equilibrium solution can be obtained through simulated game theory, enabling it to make reasonable lane-changing decisions and thus cope with scenarios where autonomous vehicles and human-driven vehicles coexist on the road.
[0166] Optionally, when the second vehicle supports communication interaction, the acceleration gain of the second vehicle is calculated based on the information of the intention sequence of the second vehicle at the next n time moments. The decision sequence information of the first vehicle at the next n time moments is selected and sent to the second vehicle. The acceleration gain of the second vehicle is updated based on the intention sequence information sent multiple times by the second vehicle. The game equilibrium solution is then obtained based on the updated acceleration gain. Therefore, completing lane-changing decisions by exchanging intentions with the second vehicle through communication interaction can effectively reduce collision risk and prediction uncertainty. Furthermore, based on the idea of rolling time-domain optimization, optimizing and sending the decision sequence of the next n time moments each time can reduce decision oscillations and accelerate algorithm convergence.
[0167] Optionally, the decision sequence includes strategies E1 and E2 of the first vehicle, and the intention sequence includes strategies F1 and F2 of the second vehicle. The frequency of occurrence of each strategy combination {E1, F1}, {E1, F2}, {E2, F1}, and {E2, F2} between the first and second vehicles is calculated, and the acceleration gain of the second vehicle is calculated and updated based on these frequencies. Therefore, only the decision information of the target vehicle is needed during information exchange; detailed gain information of the other vehicle is not required, reducing information transmission complexity.
[0168] This application embodiment also provides a lane change decision device, which has an acquisition module (e.g., Figure 2 The information acquisition module 31, the target lane selection module 32) and the decision-making module (e.g.) Figure 2 The system includes an irrational model prediction intention module 35 and an information interaction game module 34. The acquisition module acquires the target lane that the first vehicle intends to enter. The decision module calculates the acceleration gain of the second vehicle in the case of the first vehicle changing lanes, based on the first vehicle's first acceleration before the first vehicle expresses its lane-changing intention and the second vehicle's second acceleration after the first vehicle expresses its lane-changing intention. The second vehicle is a vehicle in the target lane and is located behind the first vehicle. Based on the second vehicle's acceleration gain, the game equilibrium solution is obtained. Based on the game equilibrium solution, the first vehicle makes a lane-changing decision. The lane-changing intention expression behavior may be, for example, activating a turn signal or laterally shifting towards the target lane within the driving lane.
[0169] As shown above, the benefit of the second vehicle is calculated based on the acceleration of the second vehicle before and after the first vehicle expresses its intention to change lanes. The calculated benefit is closer to the actual benefit of the second vehicle, which can more accurately predict the behavior of the second vehicle, make a more reasonable lane-changing decision for the first vehicle, and reduce the collision risk caused by prediction uncertainty.
[0170] Optionally, the decision module calculates an aggression value based on the first acceleration and the second acceleration. The aggression value indicates the aggression of the second vehicle; the higher the aggression value, the lower the benefit of the second vehicle. The aggression value when both the first acceleration and the second acceleration are above a preset value is greater than or equal to the aggression value when either the first acceleration or the second acceleration is less than a preset value.
[0171] Optionally, the decision module calculates the acceleration gain of the second vehicle based on the first acceleration, the second acceleration, and the collision probability between the first vehicle and the second vehicle. The smaller the collision probability, the greater the absolute value of the acceleration gain of the second vehicle.
[0172] Optionally, when the second vehicle does not support communication interaction, the decision module calculates the acceleration gain of the second vehicle in the case of the first vehicle changing lanes based on the first acceleration and the second acceleration, and obtains the game equilibrium solution based on the acceleration gain. When the second vehicle supports communication interaction, the decision module calculates the acceleration gain of the second vehicle based on the information of the intention sequence of the second vehicle at n future time moments, selects the information of the decision sequence of the first vehicle at n future time moments, and sends it to the second vehicle. The decision module updates the acceleration gain of the second vehicle based on the information of the intention sequence sent by the second vehicle multiple times, and obtains the game equilibrium solution based on the updated acceleration gain.
[0173] Optionally, the decision sequence includes strategies E1 and E2 of the first vehicle, and the intention sequence includes strategies F1 and F2 of the second vehicle. The decision module calculates the number of occurrences of each strategy combination {E1, F1}, {E1, F2}, {E2, F1}, and {E2, F2} of the first vehicle and the second vehicle, and calculates and updates the acceleration gain of the second vehicle based on the number of occurrences.
[0174] This application also provides a vehicle having the above-described lane change decision device.
[0175] This application also provides a computing device, including: at least one processor; and at least one memory storing program instructions, wherein the at least one processor executes the program instructions to perform the above-described lane change decision method or to perform the functions of the above-described lane change decision device.
[0176] This application also provides a computer-readable storage medium storing program instructions thereon, which a computer executes to perform the above-described lane change decision method or functions as the above-described lane change decision device.
[0177] The interactive decision-making method provided in this application can be used not only for automatic lane-changing decisions but also for decision support. Specifically, it can combine the current traffic environment of the entire road to optimize the strategies of all vehicles on the road from a global perspective, aiming to maximize road utilization efficiency and ensure safety, and then send these strategies to the vehicles for reference. Furthermore, it can be applied to platooning, where a vehicle in a platoon enters or leaves the platoon. Through interactive decision-making using game theory, it determines when to enter or leave, and at what speed and acceleration to leave, thereby minimizing interference with other vehicles in the platoon.
[0178] The embodiments of this application have been described above, but this application is not limited to the above embodiments. Various modifications can be made without departing from the spirit of this application.
Claims
1. A lane change decision method, characterized by, The method comprises: obtaining a target lane into which a first vehicle intends to enter; calculating, according to a first acceleration and a second acceleration of a second vehicle, an aggression value of the second vehicle in a situation of lane changing of the first vehicle, the second vehicle being a vehicle on the target lane and located behind the first vehicle, the first acceleration being an acceleration of the second vehicle before the first vehicle performs a lane changing intention indicating behavior, and the second acceleration being an acceleration of the second vehicle after the first vehicle performs the lane changing intention indicating behavior; calculating, according to the aggression value, an acceleration benefit of the second vehicle in the situation of lane changing of the first vehicle, wherein the greater the aggression value is, the smaller the acceleration benefit of the second vehicle is; obtaining a game equilibrium solution according to the acceleration benefit of the second vehicle; making a lane changing decision of the first vehicle according to the game equilibrium solution. In the method, the calculating, according to the first acceleration and the second acceleration of the second vehicle, of the aggression value of the second vehicle in the situation of lane changing of the first vehicle comprises: determining a first aggression value in a case where the first acceleration and the second acceleration satisfy a first condition, the first aggression value maximizing a difference between the acceleration benefits of the second vehicle before and after the first vehicle performs the lane changing intention indicating behavior, the first condition comprising that the first acceleration is less than 0 or the second acceleration is less than 0; and determining a second aggression value in a case where the first acceleration and the second acceleration do not satisfy the first condition, the second aggression value being a constant greater than 0, the second aggression value being greater than or equal to the first aggression value.
2. The method according to claim 1, wherein the aggression value in a case where the first acceleration and the second acceleration are both greater than a preset value is greater than or equal to the aggression value in a case where one of the first acceleration and the second acceleration is less than the preset value.
3. The method according to claim 1 or 2, wherein the acceleration benefit of the second vehicle is calculated according to the first acceleration, the second acceleration, and a collision probability of the first vehicle and the second vehicle, the smaller the collision probability is, the greater the absolute value of the acceleration benefit of the second vehicle is.
4. The method according to any one of claims 1-3, wherein in a case where the second vehicle does not support communication interaction, the acceleration benefit of the second vehicle in the situation of lane changing of the first vehicle is calculated according to the first acceleration and the second acceleration, and the game equilibrium solution is obtained according to the acceleration benefit of the second vehicle.
5. The method according to any one of claims 1-4, wherein the acceleration benefit of the second vehicle is calculated using a benefit function U(a) as follows:
6. The method according to claim 4, wherein wherein a is an acceleration of the second vehicle, , As shown in the following formula: wherein is a constant, is a probability of a collision of the second vehicle with the first vehicle corresponding to the acceleration a, as follows: wherein τ is a prediction time length, Δd is a longitudinal distance between the first vehicle and the second vehicle, is an estimate of the speed of the first vehicle, subject to a Gaussian distribution with mean 0 and variance σ, is a current speed of the second vehicle, As shown in the following formula: wherein is a constant, is the aggression value, as shown in the following equation: where t is a time before the first vehicle performs a lane change intention indicating behavior, a(t) is the first acceleration, t+1 is a time after the first vehicle performs the lane change intention indicating behavior, a(t+1) is the second acceleration, is a constant greater than 0. When the second vehicle supports communication interaction, according to the information of the intention sequence of the second vehicle at the next n time points, the acceleration benefit of the second vehicle is calculated, the information of the decision sequence of the first vehicle at the next n time points is selected, and is sent to the second vehicle, According to the information of the intention sequence sent by the second vehicle multiple times, the acceleration benefit of the second vehicle is updated, and the game equilibrium solution is obtained according to the updated acceleration benefit.
7. The lane-changing decision method according to claim 6, wherein the decision sequence comprises strategies E1 and E2 of the first vehicle, and the intention sequence comprises strategies F1 and F2 of the second vehicle, The number of occurrences of each of the strategy combinations {E1, F1}, {E1, F2}, {E2, F1}, and {E2, F2} of the first vehicle and the second vehicle is calculated, and the acceleration benefit of the second vehicle is calculated and updated according to the number of occurrences.
8. The lane-changing decision method according to any one of claims 1-7, wherein the lane-changing intention represents a behavior of turning on a turn signal or laterally deviating in a target lane.
9. A lane-changing decision device, comprising: an acquisition module and a decision module, The acquisition module is configured to acquire a target lane intended to be entered by a first vehicle, The decision module is configured to: According to a first acceleration and a second acceleration of a second vehicle, calculate an aggression value of the second vehicle in the case of lane-changing of the first vehicle, the aggression value indicating the aggressiveness of the second vehicle, the second vehicle being a vehicle on the target lane and located behind the first vehicle, the first acceleration being an acceleration of the second vehicle before the first vehicle implements a lane-changing intention representing behavior, and the second acceleration being an acceleration of the second vehicle after the first vehicle implements the lane-changing intention representing behavior; According to the aggression value, calculate an acceleration benefit of the second vehicle in the case of lane-changing of the first vehicle, wherein the greater the aggression value, the smaller the acceleration benefit of the second vehicle; According to the acceleration benefit of the second vehicle, obtain a game equilibrium solution; According to the game equilibrium solution, make a lane-changing decision of the first vehicle; The decision module is further configured to: in the case that the first acceleration and the second acceleration satisfy a first condition, determine a first aggression value, the first aggression value maximizing the difference between the acceleration benefits of the second vehicle before and after the first vehicle implements the lane-changing intention representing behavior, the first condition including that the first acceleration is less than 0 or the second acceleration is less than 0; in the case that the first acceleration and the second acceleration do not satisfy the first condition, determine a second aggression value, the second aggression value being a constant greater than 0, the second aggression value being greater than or equal to the first aggression value.
10. The lane-changing decision device according to claim 9, wherein The aggressive value when both the first acceleration and the second acceleration are above the preset value is greater than or equal to the aggressive value when one of the first acceleration and the second acceleration is less than the preset value.
11. The lane-changing decision device according to claim 9 or 10, wherein The decision module calculates the acceleration benefit of the second vehicle according to the first acceleration, the second acceleration and the collision probability of the first vehicle and the second vehicle, and the smaller the collision probability is, the greater the absolute value of the acceleration benefit of the second vehicle is.
12. The lane-changing decision device according to any one of claims 9-11, wherein When the second vehicle does not support communication interaction, the decision module calculates the acceleration benefit of the second vehicle in the case of lane-changing of the first vehicle according to the first acceleration and the second acceleration, and obtains the game equilibrium solution according to the acceleration benefit of the second vehicle.
13. The lane-changing decision device according to any one of claims 9-12, wherein The decision module calculates the acceleration benefit of the second vehicle using the following benefit function U(a): wherein a is an acceleration of the second vehicle, , As shown in the following formula: wherein is a constant, is a probability of a collision of the second vehicle with the first vehicle corresponding to the acceleration a, as follows: wherein τ is a prediction time length, Δd is a longitudinal distance between the first vehicle and the second vehicle, is an estimate of the speed of the first vehicle, subject to a Gaussian distribution with mean 0 and variance σ, is a current speed of the second vehicle, As shown in the following formula: wherein is a constant, is the aggression value, as shown in the following equation: where t is a time before the first vehicle performs a lane change intention indicating behavior, a(t) is the first acceleration, t+1 is a time after the first vehicle performs the lane change intention indicating behavior, a(t+1) is the second acceleration, is a constant greater than 0.
14. The lane-changing decision device according to claim 12, wherein When the second vehicle supports communication interaction, the decision module calculates the acceleration benefit of the second vehicle according to the information of the intention sequence of the second vehicle at the future n time points, selects the information of the decision sequence of the first vehicle at the future n time points and sends it to the second vehicle, The decision module updates the acceleration benefit of the second vehicle according to the information of the intention sequence sent by the second vehicle multiple times, and obtains the game equilibrium solution according to the updated acceleration benefit.
15. The lane-changing decision device according to claim 14, wherein The decision sequence includes the strategies E1 and E2 of the first vehicle, and the intention sequence includes the strategies F1 and F2 of the second vehicle, The decision module calculates the number of occurrences of each of the strategy combinations {E1, F1}, {E1, F2}, {E2, F1} and {E2, F2} of the first vehicle and the second vehicle, and calculates and updates the acceleration benefit of the second vehicle according to the number.
16. The lane-changing decision device according to any one of claims 9-15, wherein The lane-changing intention represents the behavior of turning on a turn signal or laterally deviating in the direction of the target lane within the driving lane.
17. A vehicle characterized by comprising: The lane-changing decision device according to any one of claims 9-16.
18. A computing device, comprising: Comprise: At least one processor; And At least one memory storing program instructions, The at least one processor executes the method according to any one of claims 1-8 by executing the program instructions.
19. A computer readable storage medium having stored thereon program instructions, wherein, The computer executes the method according to any one of claims 1-8 by executing the program instructions.
Citation Information
Patent Citations
Automatic driving vehicle lane changing conflict coordination model building method based on game theory
CN110362910A