Interactive perception automatic driving method based on bimodule game and deep inverse reinforcement learning
By combining Stackelberg games and Bayesian Stackelberg games with deep inverse reinforcement learning, the information asymmetry problem of multi-vehicle interaction systems in autonomous driving is solved, and a safe and efficient driving strategy is generated to adapt to complex traffic scenarios where V2V is available and unavailable.
Patent Information
- Application Number
- CN202510740142.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-08
AI Technical Summary
The existing autonomous driving decision-making and planning methods have information asymmetry and uncertainty in handling multi-vehicle interaction systems, resulting in suboptimal decision-making. Especially in highway lane change scenarios, the existing technology is difficult to effectively deal with the situation where V2V is unavailable, and the calculation complexity is high and the flexibility is poor.
Using a method based on dual-mode game and deep inverse reinforcement learning, a combination of Stackelberg game and Bayesian Stackelberg game is used to dynamically adjust the vehicle type probability to form an optimal response strategy, and combining deep inverse reinforcement learning to generate a safe and efficient driving strategy.
It improves the decision-making robustness and efficiency of autonomous driving systems in complex traffic scenarios, and can generate reasonable driving strategies when V2V is available and unavailable, improving safety and decision-making flexibility.
Smart Images

Figure CN120276261A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving, and specifically proposes a decision-making and planning method based on two-mode game and deep inverse reinforcement learning (DIRL) to achieve interactive perception planning in highway lane-changing scenarios. Background Art
[0002] Game theory is a mathematical theory that studies how multiple rational decision-makers formulate strategies in interaction. In autonomous driving, the decisions of traffic participants (ego vehicle, other vehicles, pedestrians, etc.) influence each other, forming a typical multi-agent interaction problem. Game theory can be divided into complete information games and incomplete information games according to information. Although the existing complete information game theory has important value in analyzing strategic interactions and equilibrium predictions, its theoretical framework and assumptions have some significant drawbacks in practical applications. These mainly include overly idealized assumptions, excessive reliance on the assumption of perfect rationality, the dilemma of multiple equilibria, the limitations of static analysis, computational complexity and practicality challenges, ignoring social and cultural factors, inability to explain the evolution of cooperative behavior, and insufficient handling of uncertainty. In the complex traffic scenarios of autonomous driving, the driving intentions and types of other vehicles (such as aggressiveness, cautiousness) are usually incomplete information, and in actual traffic, the reactions of other vehicles may be unpredictable. The existing game theory is applicable to simple cooperation scenarios, such as platoon following, and is not applicable to complex dynamic interactions, such as merging and emergency avoidance.
[0003] The Stackelberg game is a model in game theory, which is divided into a leader and a follower. The leader acts first, and the follower reacts according to the leader's action. The Bayesian Stackelberg game is a type of incomplete information dynamic game, which combines the sequential decision-making (leader-follower structure) of the Stackelberg game and the incomplete information of the Bayesian game (the types of participants are unknown and rely on probability beliefs). In the lane-changing scenario, the leader (autonomous vehicle) acts first, and the follower (surrounding vehicles) responds subsequently, which conforms to the reaction mode of other drivers when a vehicle changes lanes in reality and is more in line with the actual traffic dynamics. Through Bayesian updating, the AV dynamically adjusts its belief about the types of surrounding vehicles.
[0004] Deep Inverse Reinforcement Learning (DIRL) is an interdisciplinary field that combines Inverse Reinforcement Learning (IRL) and Deep Learning.
[0005] For the following published patent application with the application number CN202111053174.0 and the title "A Method for Simultaneous Jamming and Monitoring Based on Bayesian Stackelberg Game", based on the confrontation scenario model between an intelligent jammer and an enemy communication transceiver pair, using full-duplex technology, it models the communication confrontation between the friendly and enemy users under incomplete information conditions as a Bayesian Stackelberg game model, transforms the problem of simultaneously implementing jamming and monitoring into a game optimization problem, uses successive convex approximation (SCA) to transform the non-convex optimization problems of the leader and the follower, and solves the Bayesian Stackelberg game equilibrium solution through the KKT conditions. In the above patent application, the dynamic Bayesian update probability is not directly used. It only models the confrontation scenario under incomplete information through the Bayesian probability framework, which belongs to the application of Bayesian game theory rather than dynamic Bayesian inference.
[0006] For the patent application with the application number 202210596647.X and the title "A Lane-Changing Decision-Making Method for Freight Vehicles' Interactive Game Based on a Lane-Changing Decision System", aiming at the lane-changing decision-making problem of freight vehicles in a mixed traffic environment (coexisting with manual driving and autonomous driving), it proposes a decision-making method combining sensor data, vehicle-to-everything (V2X) and game theory, aiming to reduce the risk of lane-changing accidents and improve transportation efficiency. However, the situation without a V2X platform is not considered. V2X is still in development and gradually covers urban and highway scenarios.
[0007] Another example is the patent application with the application number CN202210978790 and the title "A Human-Like Decision-Making, Planning and Control Method for Autonomous Driving Considering Vehicle Interaction". It establishes a dynamic driver style recognition model to identify the driving styles of surrounding vehicles, designs a decision cost function considering traffic safety, traffic efficiency and driving comfort, establishes a lane-changing decision model considering the driving styles of both players based on the complete information non-cooperative game theory, solves the dynamic interaction behavior and decision-making between the two players in the game by introducing the Stackelberg game, completes the risk perception of the road environment by establishing a driving risk field model, and proposes a model predictive control algorithm based on the risk field. The Stackelberg game assumes complete information (vehicle states and intentions are public). If there is information asymmetry (such as the intention of the follower is unknown), it may lead to sub-optimal decisions. In the highway lane-changing scenario, the Stackelberg game is suitable for structured environments (such as closed test fields, V2X-covered sections), and the game under complete information conditions is not suitable for the complexity and uncertainty of open roads.
[0008] As described above, existing decision-making and planning methods for autonomous driving are all based on rule systems, such as finite state machines. Such methods rely on manually designed rules, have poor flexibility, are only effective in specific scenarios, have high computational complexity, require a large amount of data and training, and have obvious uncertainties when used to handle multi-vehicle interaction systems.
[0009] Currently, V2V technology has been implemented in some mass-produced vehicles and demonstration areas, but large-scale popularization still needs to solve problems such as standard unification, cost reduction, and regulation improvement. In the future, with the maturity of 5G and C-V2X, V2V will become one of the core technologies of intelligent transportation and autonomous driving. However, how to enable the coexistence of V2V vehicles and vehicles that cannot implement V2V is still a technical problem that needs to be solved urgently.
[0010] In view of this, the present application is specifically proposed. Summary of the Invention
[0011] The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning described in the present application aims to solve the problems existing in the above-mentioned prior art by using Stackelberg game and Bayesian Stackelberg game to comprehensively judge the availability of V2V, form an optimal response according to its own type, implement the overall decision-making and planning means that the leader first announces the lane-changing intention and the follower adjusts the speed according to its own driving style, so as to achieve a more practical use purpose of improving safety and efficiency.
[0012] To achieve the above design purpose, the interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning divides the decision-making into two modes: Stackelberg game and Bayesian Stackelberg game based on the integrity of information; if all relevant vehicles provide interactive intentions through V2V, the Stackelberg game is adopted; if the intentions of some vehicles are opaque, it is necessary to infer the opponent's behavior through sensors and probability models, triggering the Bayesian Stackelberg game; Specifically, it includes the following steps, Step (1) Environment perception and data collection; conduct V2V communication detection, receive the BSM of surrounding vehicles in real time, extract data, and obtain the states of surrounding vehicles through multi-sensor fusion; Step (2) V2V availability determination; Step (3) Select game mode; If V2V is available, that is, a complete information scenario, select the Stackelberg game and execute the following step (3.1); if V2V is not available, that is, an incomplete information scenario, then select the Bayesian Stackelberg game and execute the following step (3.2); Step (3.1) performs the selection of Stackelberg game; The vehicle attempting to change lanes is the leader, and the obstacle vehicle is the follower. The decision-making process is carried out in stages. The leader acts first, and the follower then optimizes; Step (3.2) performs the selection of Bayesian Stackelberg game; The autonomous vehicle acts first as the leader, and the surrounding vehicles act as followers and respond according to the actions of the leader, while considering the types of vehicles; Step (4) Deep inverse reinforcement learning planning; Learn the reward function of human drivers from natural driving data, and use the reward function to evaluate the human-likeness of candidate trajectories; Step (5) Planning; The bimodal game decision module selects the target gap. The deep inverse reinforcement learning planning module samples candidate trajectories within the target gap range. For each candidate trajectory, the FIRL is used to predict the response trajectory of the FVTG. The reward network calculates the joint reward of the candidate trajectory and the FVTG response trajectory, and selects the trajectory with the highest reward as the planning result. If the planned joint trajectory causes a conflict, the decision module is triggered to re-select the gap.
[0013] Furthermore, the said step (2) includes, Step (2.1) Communication module hardware status detection; Including signal strength detection and module activity check; Step (2.2) Communication quality index calculation; including packet loss rate and end-to-end delay; Step (2.3) Data consistency verification; including spatio-temporal alignment, error metric, and dynamic rationality verification; Step (2.4) V2V penetration rate evaluation; Step (2.5) Comprehensive availability determination; all the indicators in the above steps are weighted and fused to calculate and generate an availability score.
[0014] Furthermore, the said step (3.1) includes the following steps, Step (3.1.1) Participant division; The host vehicle is the vehicle attempting to change lanes. The host vehicle is the leader, and the obstacle vehicle is the vehicle behind in the target lane. The obstacle vehicle is the follower; Step (3.1.2) Utility function design; Consider the combined cost function of safety, traffic efficiency, and ride comfort to determine the Nash equilibrium; In the lane-changing behavior, the safety cost function of vehicle driving is manifested both horizontally and vertically; the formula of the safety cost function is as follows: (21) Wherein, and respectively represent the safety cost functions in the horizontal and vertical directions; ; Ride comfort is related to the lateral acceleration and longitudinal acceleration, and the cost function expression of comfort is as follows: (22) In the formula, and are respectively the weight coefficients of the lateral and longitudinal accelerations, and are respectively the lateral and longitudinal accelerations; The expression of the cost function of efficiency is as follows: (23) In the formula, represents the maximum vehicle speed of the lane, represents the longitudinal speed of the host vehicle, represents the longitudinal speed of the vehicle in front in the current lane, represents the relative distance between the host vehicle and the vehicle behind in the target lane, represents the minimum safety distance; The cost function of the host vehicle is expressed as: (24) In the formula, is the weight coefficient of the safety cost, is the weight coefficient of the comfort cost, is the weight coefficient of the efficiency cost; Step (3.1.3) Lane change decision; In the Stackelberg game, when making a lane change decision, the vehicle follows the principle of minimizing cost. The follower reacts to the leader's behavior according to its own cost function and the influence of the leader's behavior; The two-vehicle game optimization problem is expressed as: (25) (26) (27) (28) In the formula, represents the possible acceleration, represents the optimal acceleration of the host vehicle, represents if vehicle 1 is changing lanes, represents if it is optimal for vehicle 1 to change lanes, , Represents the behavioral decision-making of the vehicle, , represents the total cost function of the vehicle, represents the optimal decision of vehicle 2 under the influence of vehicle 1, , is the set of possible actions of the vehicle; Step (3.1.4) outputs the target gap to the deep inverse reinforcement learning planning module.
[0015] Furthermore, step (3.2) performs the selection of the Bayesian Stackelberg game; The autonomous vehicle acts first as the leader, and the surrounding vehicles act as followers and respond according to the actions of the leader, while considering the types of vehicles; specifically, it includes the following steps: Step (3.2.1) defines the game roles and order; The host vehicle acts as the leader and acts first, and the vehicle behind in the target lane acts as the follower and selects a response strategy according to the actions of the leader; The host vehicle selects a strategy based on the probability distribution of the types of followers ; after observing the actions of the host vehicle, the follower selects the optimal response according to its own type ; ; Step (3.2.2) models incomplete information; It includes the type space and probability distribution and utility function; the host vehicle knows its own type, but only has a prior distribution of the type of FVTG , and FVTG knows its own type , but is not sure about the specific strategy of the leader; The utility of the host vehicle needs to consider the response strategy of the follower and its type distribution: ; FVTG selects the optimal response according to the actions of the EV and its own type: ; Step (3.2.3) designs the utility function; Designs the utility function by integrating safety, efficiency, and traffic friendliness; Step (3.2.4) solves the Bayesian Stackelberg equilibrium; Step (3.2.5) performs probability update and dynamic adjustment; The host vehicle predicts the real-time behavior of the follower through GMM, and updates the distribution of the follower type using Bayes' theorem, as shown in the following formula: (35) where, is generated by the SIDM model; Combines the R-R diagram of the time headway THW and acceleration to update the aggressiveness of the follower in real time , for adjusting type probabilities; The output of step (3.2.6) is the target gap to the deep inverse reinforcement learning planning module.
[0016] Furthermore, the said step (3.2.3) includes: Step (3.2.3.1) Safety cost function design; During a lane change, for vehicle i, its driving safety within a certain prediction range h is determined by comparing the distance between vehicles and a predefined threshold safety distance , as well as the lane positioning of each vehicle and ; If the relative lane difference and are for all , then vehicle i is safe; if the vehicle is crossing a lane boundary, the vehicle's lane can be multiple; in this case, the safety of vehicle i can only be ensured when all relevant vehicles meet the definition of driving safety; In the above definition, is the relative lane difference ; is the distance between vehicle i and vehicle j at time t, as follows: (29) In the formula, h>0 is the time interval, and are the longitudinal driving speeds of vehicle i and vehicle j at time t; Many alternative safety metrics have been used to quantify the safety level, such as time headway (THW), time to collision (TTC), and the distance between vehicles; in addition, when a dangerous situation occurs, the driver's deceleration reaction time also plays a key role; In summary, these factors are integrated into the required deceleration, as follows: (30) That is, under certain conditions, the deceleration required for the following vehicle to avoid a collision, represents the distance between vehicles, and represent the speeds of the front and rear vehicles, represents the driver's reaction time, represents the acceleration of the following vehicle, represents the maximum braking deceleration; the driving safety metric is determined as: (31) In the formula, is the weight factor; Step (3.2.3.2) Design of traffic - friendliness cost function; Traffic - friendliness is represented by restricting the lane - changing frequency; lane - changing is only initiated when the current lane cannot meet the driving function; if the required deceleration in the future is less than the threshold of the current lane then a lane - changing penalty will be effectively imposed. The traffic - friendliness factor can be expressed as: (32) In the formula, is the weight factor, is the threshold deceleration to ensure driving safety, is the predicted required deceleration calculated based on the future state of the vehicle, is the sign function; Step (3.2.3.3) Design of efficiency cost function; The driving efficiency of vehicle i is determined by the difference between the current speed and the future speed . is the vehicle speed in the long - term stable car - following state. The expression of the efficiency cost function is: (33) In the formula, represents the limit speed that vehicle i can freely accelerate to when it is greater than the safe car - following distance , is the weight factor.
[0017] Furthermore, the step (3.2.4) includes: Step (3.2.4.1) Backward induction method; Follower response modeling, for each possible leader strategy and follower type , solve the optimal response of the follower ; Leader strategy optimization, the host vehicle, as the leader, needs to maximize its own expected utility, considering the uncertainty and response of the follower type, as follows: (34) In the formula, is the leader strategy, is the follower type, is the prior distribution of the follower, is the optimal response of the follower, is the utility of the host vehicle; Step (3.2.4.2) Bilevel optimization problem; Upper - level problem, leader optimizes the strategy ; Lower - level problem, for each , follower solution ; Solve the double-layer optimization problem using SQL.
[0018] Furthermore, the step (4) includes Step (4.1) Data preparation and preprocessing; Use the NGSIM public driving dataset to screen expert trajectory segments containing highway lane-changing scenarios; construct a dataset suitable for DIRL training to ensure unified input format and complete features; Step (4.2) Adopt sampling-based deep inverse reinforcement learning; According to the sampling-based DIRL to solve the problem of the solution difficulty of DIRL. According to the principle of maximum entropy, the probability of selecting a trajectory is proportional to the natural exponent of its reward; Step (4.3) Deep inverse reinforcement learning planning; Select the trajectory with the highest reward value through sampling, evaluation, and selection.
[0019] Furthermore, the step (4.2) solves the problem of the solution difficulty of DIRL according to the principle of maximum entropy. The probability of selecting a trajectory is proportional to the natural exponent of its reward, as shown in the following formula: (36) Where represents the probability that the trajectory τ is selected, is the reward of the trajectory τ, is the parameter of the reward function, is called the normalization function or partition function, and D is the set of all possible trajectories that the agent can choose; The continuous and huge state space makes it difficult to calculate the partition function in the lane-changing scenario. The following partition function can be adopted to approximately represent it by sampling candidate trajectories in the state space: (37) Where represents all sampled candidate trajectories, is the trajectory 's reward; then the probability that the trajectory τ is selected is: (38) Where All the trajectories in have the same initial state as the trajectory τ; Adopt maximum entropy inverse reinforcement learning to maximize the likelihood of the expert demonstration trajectory by adjusting the parameter of the reward function, as shown in the following formula: (39) Among them, $E$ represents all expert demonstration trajectories; then the objective function of IRL is given by the following formula: (40) where $E$ is all expert demonstration trajectories, is a single expert demonstration trajectory, is the trajectory 's reward,
[0020] is the trajectory 's probability of being selected; The gradient of the derived reward function parameters is: (41) In the formula, is the trajectory 's probability of being selected.
[0021] Furthermore, the said step (4.3) includes: Step (4.3.1) Candidate trajectory sampling; It includes a longitudinal adjustment stage and a lateral lane change stage; the longitudinal adjustment stage samples high-dimensional parameters within the time interval , the target longitudinal position , speed , adjustment time ; Use the following fifth-degree polynomial to fit the first-stage longitudinal position of the vehicle in the S-L coordinate system: (42) where, is the fifth-degree polynomial coefficient, is the fifth-degree polynomial coefficient, is the fifth-degree polynomial coefficient, is the fifth-degree polynomial coefficient, is the fifth-degree polynomial coefficient, is the fifth-degree polynomial coefficient, and $t$ is the time variable; Then the speed and acceleration of the vehicle are: (43) (44) The boundary conditions include the initial state and the target state , then the coefficients of the polynomial can be expressed as: (45) where, is the initial time, is the target time, is the initial longitudinal position, is the initial speed, is the initial acceleration, is the target longitudinal position, is the target speed, is the target acceleration; In the lateral lane change phase, the lateral trajectory of the lane change is fitted by formula (45); the L coordinate in the S-L coordinate system is expressed as: (46) where, are the coefficients of the fifth-degree polynomial, are the coefficients of the fifth-degree polynomial, are the coefficients of the fifth-degree polynomial, are the coefficients of the fifth-degree polynomial, are the coefficients of the fifth-degree polynomial, are the coefficients of the fifth-degree polynomial; The candidate trajectory can be generated by sampling the high-level driving intention The required time The sampling range of is The sampling range of the target speed is is the speed of the vehicle behind the ego vehicle in the target lane and the vehicle closest to the ego vehicle in the vehicle behind, the target longitudinal position The sampling range of m, the sampling interval is 2m, where is The position of the vehicle behind the ego vehicle in the target lane at time is The position of the vehicle in front of the ego vehicle in the target lane at time; Any candidate trajectory that causes a collision will be removed; Step (4.3.2) Interactive perception prediction; In the lane change scenario, the ego vehicle mainly affects the vehicle behind the target gap. The vehicle behind the target gap is represented by FVTG. The trajectory of FVTG is predicted using feature-based inverse reinforcement learning; FIRL learns the internal reward function of FVTG from driving data and finds the trajectory that maximizes the reward function as the predicted trajectory; The lane change scenario is regarded as a car-following scenario where the leading vehicle in FVTG suddenly changes; FVTG can recognize the lane change intention of the ego vehicle and regard it as a new leading vehicle when the ego vehicle uses its turn signal or historical trajectory to send a lane change signal; In this scenario, the feature selection of FVTG is as follows: 1) Acceleration: (47) where, acc is the instantaneous acceleration of FVTG; 2) Relative speed with respect to the ego vehicle: (48) Wherein, is the relative speed between the FVTG and the ego vehicle; 3) Relative distance with respect to the ego vehicle: (49) Wherein, d is the actual relative distance between the FVTG and the ego vehicle, represents the expected relative distance to the ego vehicle; the distance between the FVTG and the LVTG is regarded as the desired relative distance; 4) Collision penalty: (50) In feature-based inverse reinforcement learning, the reward function is as follows: (51) Wherein is the weight parameter of the acceleration feature, is the weight parameter of the relative speed feature, is the weight parameter of the relative distance feature, is the weight parameter of the collision penalty feature; Given the reward function, the optimal trajectory of the FVTG is planned using MPC; that is, the reward function of the traffic vehicle FVTG is learned through FIRL, and its trajectory is generated through model predictive control MPC; Step (4.3.3) Reward network evaluation; A reward network is used to represent the reward function used by humans for planning, and the reward network is trained using sampling-based DIRL; Step (4.3.4) Sort the candidate trajectories according to the reward values, and select the trajectory with the highest reward value as the final planning result.
[0022] Furthermore, the said step (5) includes, Step (5.1) Conflict detection; Step (5.2) Reward evaluation and threshold setting; The reward network will output a comprehensive score for each candidate trajectory, including safety, efficiency, comfort, and interaction cost; safety threshold, if the score is lower than the safety threshold, it is determined as an infeasible trajectory; dynamic adjustment, the threshold can be dynamically adjusted according to the scenario, for example, increasing the safety requirements in dense traffic; Step (5.3) Logic to trigger re-decision; When a conflict or insufficient reward score is detected, the following process is triggered for adjustment: Conflict marking, the planning module marks the conflict trajectory as "high risk" and records the conflict type; The decision-making module receives conflict information, re-evaluates the feasibility of the current target gap, re-selects the gap, and adjusts the driving strategy; Update the candidate trajectory set, re-sample candidate trajectories according to the new target gap, and repeat the planning process.
[0023] In summary, the interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning proposed in this application has the following advantages and beneficial effects: 1. The dynamic interaction modeling based on the reward learning ability of bimodal game and DIRL in this application can generate safer, more efficient and more human-like driving strategies in the lane-changing scenario, which is more conducive to the development of the decision-making and planning module of L3 and above autonomous driving systems.
[0024] 2. This application improves on the existing autonomous driving decision-making and planning system and further combines bimodal game and deep inverse reinforcement learning. Based on the integrity of information, the decision-making is divided into two modes: Stackelberg game and Bayesian Stackelberg game. In the case of V2V, vehicles can obtain the real-time states of other vehicles (such as speed, acceleration, intention) through direct communication, forming an information-complete interaction environment. At this time, the Stackelberg game is adopted to optimize the collaborative decision-making using the clear leader-follower relationship. The collaborative efficiency is high. The host vehicle, as the leader, can actively plan the optimal lane-changing strategy, and other vehicles, as the followers, respond based on the clear information, reducing game conflicts. The computational efficiency is high. When the information is complete, there is no need to process complex probability distributions, and the model can be solved faster, which is suitable for scenarios with high real-time requirements. In the case of no V2V, vehicles need to rely on sensors to indirectly perceive the surrounding environment, and there is uncertainty about the intentions or types of other vehicles (such as aggressive / conservative). The Bayesian Stackelberg game processes the problem of incomplete information through probability modeling, improving the robustness of decision-making. The robustness is strong. By modeling the possible types of other vehicles through probability distributions (such as the probability of aggressive driving), the expected return in the worst case is optimized, reducing risks. It can flexibly handle uncertainties, dynamically update the belief about the behavior of other vehicles (such as based on sensor observations), gradually approach the real intention, and improve the reliability of decision-making.
[0025] 3. This application establishes a decision-making and planning system based on bimodal game and deep inverse reinforcement learning. The decision-making module and the planning module adopt a hierarchical structure. The decision-making module outputs the optimal target gap, and the planning module generates and evaluates specific trajectories within this range based on the selected gap, and finally outputs the optimal trajectory. The trajectories generated by the planning module will in turn affect the behavior prediction of traffic vehicles, forming a closed-loop interaction. If the planning result is not ideal (such as the interaction cost is too high), the decision-making module can be triggered to re-select the gap. Brief Description of the Drawings
[0026] The present application will be further described in conjunction with the following drawings; Figure 1 It is a schematic diagram of a decision-making and planning system applying the interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning described in this application; Figure 2 It is a schematic diagram of the lane-changing process; Figure 3 It is a flowchart of the Stackelberg game; Figure 4 It is a flowchart of the Bayesian Stackelberg game; Detailed implementation manners
[0027] To further elaborate on the technical means adopted by this application to achieve the predetermined design purpose, the following relatively preferred implementation manners are proposed in combination with the accompanying drawings.
[0028] Specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementation manners disclosed below.
[0029] As Figure 1 shown, the decision-making and planning system applying the interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning described in this application includes: A bimodal game decision-making module, based on the integrity of information, divides autonomous driving decisions into two modes, namely the Stackelberg game and the Bayesian Stackelberg game, in the highway lane-changing scenario; selects the Stackelberg game when V2V is available. The host vehicle constructs a two-layer optimization problem through the Stackelberg game model, and iteratively optimizes its own strategy on the basis of predicting the response of the obstacle vehicle until the Nash equilibrium is reached; selects the Bayesian Stackelberg game when V2V is not available, and adjusts the probability of the surrounding vehicle types in real time based on Bayes' theorem to improve the decision-making adaptability; A deep inverse reinforcement learning planning module, learns the reward function of human drivers from natural driving data, and uses the reward function to evaluate the humanity of candidate trajectories; extracts the spatio-temporal features of the ego vehicle trajectory through a convolutional neural network (CNN) encoder, and a multi-layer perceptron (MLP) fusion unit encodes the joint state of the ego vehicle and other vehicles into a reward value, and uses the maximum entropy optimization framework to maximize the likelihood probability of the expert trajectory through gradient ascent; The phased collaboration module conducts decision-making and planning in blocks; first, decision-making is carried out, and after obtaining the target gap, planning is carried out; the target gap is selected based on the game result to narrow the planning space, and the planning module generates and evaluates specific trajectories within this range based on the selected gap, and finally outputs the optimal trajectory; the trajectories generated by the planning module will in turn affect the behavior prediction of traffic vehicles, forming a closed-loop interaction; if the planning result is not ideal, the decision-making module can be triggered to reselect the gap. Among them, both the dual-mode game decision-making module and the deep inverse reinforcement learning planning module adopt a hierarchical structure. The dual-mode game decision-making module outputs the target gap to the deep inverse reinforcement learning planning module, including the initial state and the target state; the deep inverse reinforcement learning planning module uses deep inverse reinforcement learning to select trajectories under the constraint of the target gap. If there is no suitable trajectory, it will trigger the dual-mode game decision-making module again to reselect the gap, and jointly complete the decision-making and planning of autonomous driving.
[0030] Such as Figures 2 to 4 As shown, the interactive perception autonomous driving method based on dual-mode game and deep inverse reinforcement learning described in this application divides decision-making into two modes: Stackelberg game and Bayesian Stackelberg game based on the integrity of information; if all relevant vehicles provide real-time and reliable interaction intentions (such as acceleration, lane change request) through V2V, the Stackelberg game is adopted; if there are some vehicles with opaque intentions (such as no V2V or poor communication quality), it is necessary to infer the opponent's behavior through sensors and probability models, triggering the Bayesian Stackelberg game. Specifically, it includes the following steps: Step (1) Environment perception and data collection; Conduct V2V communication detection, receive the BSM (Basic Safety Message) of surrounding vehicles in real time, and extract data including but not limited to position, speed, acceleration, and intention (lane change request); obtain the states of surrounding vehicles through multi-sensor fusion (such as through lidar, camera, millimeter wave radar, etc.). Step (2) V2V availability determination; Quantitatively evaluate from four dimensions: hardware status, data quality, communication performance, and scenario penetration rate. Step (2.1) Communication module hardware status detection; Including signal strength detection and module activity check; RSSI (Received Signal Strength Indication) reflects the physical strength of the received signal and is related to the communication distance and environmental interference; the signal strength detection judges the physical connection state of the V2V module through the received signal strength indication (RSSI), such as the following expression: (1) Among them, is the transmission power, is the transmitting antenna gain, is the receiving antenna gain, d is the vehicle spacing, and λ is the signal wavelength, is the obstacle attenuation; (2) From the above formula, if the signal strength is lower than the threshold, it is determined that the hardware connection is unstable; For the module activity check, it is to count the message sending frequency of the V2V module within the time window T, as shown in the following formula: (3) (4) Among them, is the number of received BSMs (Basic Safety Messages), and T is the time window; Step (2.2) Communication quality index calculation; It includes the packet loss rate and the end-to-end delay; The packet loss rate (Packet Loss Rate, PLR) reflects the stability of the communication link. A high packet loss rate may lead to information loss. Calculate the proportion of lost messages within the time window T according to the following formula: (5) (6) Among them, is the number of received BSMs (Basic Safety Messages); Calculate whether the packet loss rate meets the standard according to the following formula: (7) If it exceeds the above value, it is determined that the communication is unstable; The end-to-end delay (Latency) directly affects the real-time decision-making and needs to be controlled at the millisecond level; Measure the time difference from sending to receiving of the message according to the following formula: (8) Among them, is the message reception timestamp, is the message sending timestamp; Calculate whether the end-to-end delay meets the standard according to the following formula: (9) If it times out, it is determined that the communication delay is unacceptable.
[0031] Step (2.3) Data consistency verification; It includes spatio-temporal alignment, error measurement, and dynamic rationality verification; For spatio-temporal alignment, V2V data needs to be fused with local sensor (radar, camera) data to exclude false signals; Use Kalman filtering to align V2V data with local sensor (radar, camera) data as follows: (10) where, is the observation vector of V2V, is the observation vector of the sensor; K is the Kalman gain matrix to dynamically adjust the trust weight; Error metric, calculate the difference between V2V data and the fusion result; calculate the position error (Euclidean distance) as follows: (11) (12) where, is the abscissa of the vehicle reported by V2V, is the ordinate of the vehicle reported by V2V, is the abscissa of the vehicle after sensor fusion, is the ordinate of the vehicle after sensor fusion.
[0032] The speed error is calculated as follows: (13) (14) where, is the vehicle speed reported by V2V, is the vehicle speed after sensor fusion; Dynamic rationality verification, check whether the acceleration reported by V2V is within the physically feasible range as follows: (15) where, is the vehicle acceleration reported by V2V, is the maximum acceleration allowed by vehicle dynamics; Calculate whether the maximum acceleration exceeds the limit as follows: (16) If the maximum acceleration exceeds the vehicle dynamics limit, it is determined as abnormal data; Step (2.4) V2V penetration rate assessment; Statistically calculate the proportion of vehicles supporting V2V among surrounding vehicles as follows: (17) where, is the number of surrounding vehicles supporting V2V, is the total number of vehicles detected by the sensor; Calculate whether the penetration rate meets the standard according to the following formula: (18) Key vehicle screening, only count the vehicles related to lane change, including the vehicles in front of and behind the target lane and the vehicle behind in the current lane; Step (2.5) Comprehensive availability determination; Fuse all the indicators in the above steps with weights, calculate and generate an availability score according to the following formula: (19) Among them, is the weight coefficient of the hardware state, is the weight coefficient of the packet loss rate, is the weight coefficient of the communication delay, is the weight coefficient of the position error, is the weight coefficient of the V2V penetration rate; is the hardware state indication function, which is 1 if it meets the standard, otherwise 0; The final condition for determining V2V availability is as follows: (20) Step (3) Select the game mode; If V2V is available, it is a complete information scenario, select the Stackelberg game, and execute the following step (3.1); If V2V is not available, it is an incomplete information scenario, then select the Bayesian Stackelberg game, and execute the following step (3.2); Step (3.1) Execute the selected Stackelberg game; The vehicle attempting to change lanes is the leader, and the obstacle vehicle is the follower. The decision-making process is carried out in stages. The leader acts first, and the follower then optimizes; specifically, it includes the following steps: Step (3.1.1) Participant division; The host vehicle is the vehicle attempting to change lanes. The host vehicle is the leader (Leader), and the obstacle vehicle is the vehicle behind in the target lane. The obstacle vehicle is the follower (Follower); Step (3.1.2) Utility function design; Consider the combined cost function of safety, traffic efficiency, and ride comfort to determine the Nash equilibrium; In the lane change behavior, the safety cost function of vehicle driving is manifested both horizontally and vertically; the formula of the safety cost function is as follows: (21) Among them, represents the lateral safety cost function, represents the longitudinal safety cost function; ; Ride comfort is related to lateral acceleration and longitudinal acceleration, and the cost function expression of comfort is as follows: (22) In the formula, is the weight coefficient of lateral acceleration, is the weight coefficient of longitudinal acceleration, is the lateral acceleration, is the longitudinal acceleration; The expression of the cost function of efficiency is as follows: (23) In the formula, represents the maximum vehicle speed of the lane, represents the longitudinal speed of the host vehicle, represents the longitudinal speed of the vehicle in front in the current lane, represents the relative distance between the host vehicle and the vehicle behind in the target lane, represents the minimum safety distance; The cost function of the host vehicle is expressed as: (24) In the formula, is the weight coefficient of the safety cost, is the weight coefficient of the comfort cost, is the weight coefficient of the efficiency cost; Step (3.1.3) Lane change decision; In the Stackelberg game, when making a lane change decision, the vehicle follows the principle of minimizing cost. The follower reacts to the leader's behavior according to its own cost function and the influence of the leader's behavior; The two-vehicle game optimization problem is expressed as: (25) (26) (27) (28) In the formula, represents the optimal strategy of the host vehicle, represents the optimal acceleration of the host vehicle, represents the optimal lane change decision of the host vehicle, represents the possible acceleration of the host vehicle, represents whether vehicle 1 is changing lanes, Represents the acceleration of the vehicle, Indicates whether lane changing is optimal for Vehicle 1, , Represents the behavior decision of Vehicle i, , Represents the total cost function of Vehicle i, Represents the optimal decision of Vehicle 2 under the influence of Vehicle 1, , Is the set of possible actions of Vehicle i, Is the possible strategy of the ego vehicle, , Represents the longitudinal speed of Vehicle i, Represents the maximum allowed longitudinal speed, , Represents the longitudinal acceleration of Vehicle i, Represents the minimum longitudinal acceleration, Represents the maximum longitudinal acceleration; Step (3.1.4) outputs the target gap to the deep inverse reinforcement learning planning module; Step (3.2) performs the selective Bayesian Stackelberg game; The autonomous vehicle acts first as the leader, and the surrounding vehicles act as followers and respond according to the leader's actions, while considering their types (aggressive, mild, cautious); specifically, it includes the following steps: Step (3.2.1) defines the game roles and order; The host vehicle (EV) acts as the leader and takes precedence (such as a lane-changing intention), and the vehicle behind in the target lane (FVTG) acts as the follower and selects a response strategy according to the leader's action; The host vehicle selects a strategy based on the probability distribution of the follower's type (such as lane changing or keeping the lane); after observing the host vehicle's action, the follower selects the optimal response according to its own type ; ; Step (3.2.2) models incomplete information; Including the type space, probability distribution, and utility function; the host vehicle knows its own type, but only has a prior distribution of the type of FVTG (aggressive, mild, cautious) , and FVTG knows its own type , but is unsure of the specific strategy of the leader; The utility of the host vehicle needs to consider the response strategy of the follower and its type distribution: ; FVTG selects the optimal response according to the action of EV and its own type: ; Step (3.2.3) designs the utility function; Design the utility function by integrating safety, efficiency, and traffic friendliness; Step (3.2.3.1) Design the safety cost function; During the lane-changing process, for vehicle i, its driving safety within a certain prediction range h is determined by comparing the distance between vehicles and a predefined threshold safety distance , as well as the lane positioning of each vehicle and ; If the relative lane difference and are such that for all , then vehicle i is safe; if the vehicle is crossing the lane boundary, the vehicle's lane can be multiple; in this case, the safety of vehicle i can be ensured only when all relevant vehicles meet the definition of driving safety; In the above definition, is the relative lane difference ; is the distance between vehicle i and vehicle j at time t, as follows: (29) where h > 0 is the time interval, is the longitudinal driving speed of vehicle i at time t, is the longitudinal driving speed of vehicle j at time t; Many alternative safety metrics have been used to quantify the safety level, such as time headway (THW), time to collision (TTC), and the distance between vehicles; in addition, the driver's deceleration reaction time also plays a key role when a dangerous situation occurs; In summary, these factors are integrated into the required deceleration as follows: (30) That is, under certain conditions, the deceleration required for the following vehicle to avoid a collision , represents the distance between vehicles, represents the speed of the leading vehicle, represents the speed of the following vehicle, represents the driver's reaction time, represents the acceleration of the following vehicle, represents the maximum braking deceleration; therefore, the driving safety metric is determined as: (31) where is the weight factor; Step (3.2.3.2) Design the traffic friendliness cost function; Traffic friendliness is represented by restricting the lane-changing frequency; lane changes are only initiated when the current lane cannot meet the driving function; if the required deceleration in the future is less than the threshold of the current lane , then a lane-changing penalty will be effectively imposed. The traffic friendliness factor can be expressed as: (32) In the formula, is the weight factor, is the threshold deceleration to ensure driving safety, is the predicted required deceleration calculated based on the future state of the vehicle, is the sign function; Step (3.2.3.3) Efficiency cost function design; The driving efficiency of vehicle i is determined by the difference between the current speed and the future speed . is the vehicle speed in the long-term stable car-following state. The expression of the efficiency cost function is: (33) In the formula, represents the maximum speed that vehicle i can freely accelerate to when it is greater than the safe car-following distance , is the weight factor; Step (3.2.4) Bayesian Stackelberg equilibrium solution; It includes the following steps, Step (3.2.4.1) Backward induction method; Follower response modeling. For each possible leader strategy and follower type , solve the optimal response of the follower; Leader strategy optimization. The host vehicle, as the leader, needs to maximize its own expected utility, considering the uncertainty and response of the follower type, as shown in the following formula: (34) In the formula, is the leader strategy, is the follower type, is the prior distribution of the follower, is the optimal response of the follower, is the utility of the host vehicle; Step (3.2.4.2) Two-layer optimization problem; Upper-level problem, leader optimization strategy ; Lower-level problem, for each , the follower solves ; Solve the bilevel optimization problem using SQL (Sequential Quadratic Programming); Step (3.2.5) Probability update and dynamic adjustment; The host vehicle predicts the real-time behavior of the follower (such as acceleration and vehicle distance) through GMM, and updates the distribution of the follower type using Bayes' theorem, as shown in the following formula: (35) where Generated by the SIDM model; Combined with the R-R diagram of the time headway THW and acceleration, the aggressiveness of the follower is updated in real time , which is used to adjust the type probability; The R-R diagram is a two-dimensional chart based on Risk and Response, used to quantify the aggressiveness of the driver; the horizontal axis (Risk) is the risk indicator, using the time headway (THW), which represents the time interval between the current vehicle and the vehicle in front (unit: second); the vertical axis (Response) is the response indicator, using the longitudinal acceleration, which reflects the dynamic behavior of the vehicle; the slope of the R-R diagram represents the parameter that quantifies the driver's aggressiveness; Through the R-R diagram and combined with the time headway THW and acceleration, the aggressiveness of the driver can be dynamically quantified, thus providing real-time type probability update for the Bayesian Stackelberg game; Step (3.2.6) Output the target gap to the deep inverse reinforcement learning planning module; Step (4) Deep inverse reinforcement learning planning; Learn the reward function of human drivers from natural driving data, and use the reward function to evaluate the humanity of candidate trajectories; including the following steps, Step (4.1) Data preparation and preprocessing; Use the NGSIM public driving dataset to screen the expert trajectory segments containing highway lane-changing scenarios; construct a dataset suitable for DIRL training to ensure that the input format is unified and the features are complete.
[0033] Step (4.2) Adopt sampling-based deep inverse reinforcement learning; According to the principle of maximum entropy, solve the problem of the solution difficulty of DIRL, and the probability of selecting a trajectory is proportional to the natural exponent of its reward, as shown in the following formula: (36) where represents the probability that the trajectory τ is selected, is the reward of the trajectory τ, is a parameter of the reward function, which is called the normalization function or partition function, and D is the set of all possible trajectories that the agent can choose; The continuous and large state space makes it difficult to calculate the partition function in the lane-changing scenario. The following partition function can be adopted to approximately represent it by sampling candidate trajectories in the state space: (37) where, represents all sampled candidate trajectories, is the trajectory 's reward; then the probability that the trajectory τ is selected is: (38) where, all the trajectories in have the same initial state as the trajectory τ; Adopt maximum entropy inverse reinforcement learning to maximize the likelihood of the expert demonstration trajectory by adjusting the parameter of the reward function, as follows: (39) where, E represents all expert demonstration trajectories; then the objective function of IRL is given by: (40) where, E is all expert demonstration trajectories, is a single expert demonstration trajectory, is the reward of the trajectory ;
[0034] is the probability that the trajectory is selected; The gradient of the derived reward function parameter is: (41) In the formula, is the probability that the trajectory is selected; Step (4.3) Deep inverse reinforcement learning planning; Select the trajectory with the highest reward value through sampling, evaluation, and selection, which specifically includes the following steps: Step (4.3.1) Candidate trajectory sampling; Includes the longitudinal adjustment stage and the lateral lane-changing stage; The longitudinal adjustment stage samples high-dimensional parameters within the time interval , the target longitudinal position , speed , adjustment time ; Use the following fifth-degree polynomial to fit the first-stage longitudinal position of the car in the S-L coordinate system: (42) Among them, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, and t is the time variable; Then the speed and acceleration of this vehicle are: (43) (44) The boundary conditions include the initial state and the target state , then the coefficients of the polynomial can be expressed as: (45) Among them, is the initial moment, is the target moment, is the initial longitudinal position, is the initial speed, is the initial acceleration, is the target longitudinal position, is the target speed, is the target acceleration; In the lateral lane-changing stage, the above fifth-degree polynomial is used to fit the lateral trajectory of the lane change; The L coordinate in the S-L coordinate system is expressed as: (46) Among them, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial.
[0035] The candidate trajectory can be generated by sampling the high-level driving intention , and the sampling range of the required time is , the sampling interval is 1 s, and the sampling range of the target speed is , the sampling interval is 0.5 m / s, where is the speed of the vehicle behind the ego vehicle in the target lane and the vehicle closest to the ego vehicle among the vehicles behind, and the sampling range of the target longitudinal position is m, and the sampling interval is 2m, where is the position of the vehicle behind the ego-vehicle in the target lane at time is the position of the vehicle in front of the ego-vehicle in the target lane at time. Any candidate trajectory that leads to a collision will be removed; Step (4.3.2) Interactive perception prediction; In the lane change scenario, the ego-vehicle mainly affects the vehicle behind the target gap. The vehicle behind the target gap is represented by FVTG. The Feature-based Inverse Reinforcement Learning (FIRL) is used to predict the trajectory of FVTG. FIRL learns the internal reward function of FVTG from driving data and finds the trajectory that maximizes the reward function as the predicted trajectory. The lane change scenario is regarded as a car-following scenario where the leading vehicle in FVTG suddenly changes. FVTG can recognize the lane change intention of the ego-vehicle and regards it as the new leading vehicle when the ego-vehicle uses its turn signal or historical trajectory to send a lane change signal; In this scenario, the feature selection of FVTG is as follows: 1) Acceleration: (47) where acc is the instantaneous acceleration of FVTG.
[0036] 2) Speed relative to the ego-vehicle: (48) where is the relative speed between FVTG and the ego-vehicle.
[0037] 3) Relative distance to the ego-vehicle: (49) where d is the actual relative distance between FVTG and the ego-vehicle, represents the expected relative distance to the ego-vehicle; the distance between FVTG and LVTG (the vehicle in front of the target gap) is regarded as the desired relative distance (the original relative distance before the ego-vehicle cuts in); 4) Collision penalty: (50) In the feature-based inverse reinforcement learning, the reward function is as follows: (51) where is the weight parameter of the acceleration feature, is the weight parameter of the relative speed feature, is the weight parameter of the relative distance feature, is the weight parameter of the collision penalty feature; The training process of the reward function in the feature-based inverse reinforcement learning is as follows: Initialize the parameters of the reward function ; Calculate the average of the empirical feature vectors of all demonstrations ; Solve for the optimal trajectory τ based on the current reward function; Calculate the features of the optimal trajectory ; Calculate the gradient of the reward function ; According to the gradient Update the parameters of the reward function, where is the learning rate; Repeat from the third step until convergence.
[0038] Given a reward function, adopt MPC to plan the optimal trajectory of FVTG. That is, learn the reward function of the traffic vehicle (FVTG) through FIRL and generate its trajectory through model predictive control (MPC).
[0039] Step (4.3.3) Reward network evaluation; Adopt a reward network to represent the reward function used by humans for planning, and use sampling-based DIRL to train the reward network. Construct an end-to-end reward network that fuses the spatio-temporal state information of the ego vehicle and other vehicles. The network structure includes an input layer (inputting the ego vehicle trajectory time-series data and the instantaneous state of other vehicles), a CNN encoder (processing the ego vehicle trajectory), an MLP encoder (processing the state of other vehicles), feature fusion and reward prediction (outputting the reward value). The reward network is trained from human driving data through deep inverse reinforcement learning (DIRL), and the goal is to maximize the likelihood probability of the expert trajectory.
[0040] Step (4.3.4) Sort the candidate trajectories according to the reward values, and select the trajectory with the highest reward value as the final planning result.
[0041] Step (5) Planning; The two-mode game decision-making module selects the target gap. The deep inverse reinforcement learning planning module samples candidate trajectories within the target gap range. For each candidate trajectory, use FIRL to predict the response trajectory of FVTG (considering the impact of ADV actions). The reward network calculates the joint reward of the candidate trajectory and the FVTG response trajectory, and selects the trajectory with the highest reward as the planning result. If the planned joint trajectory causes a conflict (such as FVTG being unable to avoid), then trigger the decision-making module to re-select the gap; Specifically, it includes the following steps: Step (5.1) Conflict detection; The system quantifies the trajectory conflict risk through the following metrics, including the minimum relative distance (Min Distance): If the minimum distance in the predicted trajectories of the ADV and FVTG is lower than the safety threshold, it is regarded as a high risk; Time to Collision (TTC): If the TTC is lower than the critical value (1.5 seconds), it is considered that there is a collision risk; Abrupt Acceleration: If there is an extreme deceleration (the absolute value of acceleration exceeds 3 m / s²) in the predicted trajectory of the FVTG, indicating that it is forced to make an emergency avoidance, the ADV trajectory needs to be adjusted; Interaction Cost: The interaction impact is quantified by the output value of the reward function learned by FIRL. If the interaction cost (speed loss, reduced comfort) exceeds the preset threshold, re-planning is triggered; Step (5.2) Reward evaluation and threshold setting; The reward network outputs a comprehensive score for each candidate trajectory, including safety, efficiency, comfort, and interaction cost; Safety Threshold: If the score is lower than the safety threshold, the trajectory is determined to be infeasible; Dynamic Adjustment: The threshold can be dynamically adjusted according to the scenario. For example, the safety requirements are increased in dense traffic; Step (5.3) Logic for triggering re-decision; When a conflict or insufficient reward score is detected, the adjustment is triggered according to the following process: Conflict Marking: The planning module marks the conflicting trajectory as "high risk" and records the type of conflict (such as too close distance, insufficient TTC); The decision-making module receives the conflict information, re-evaluates the feasibility of the current target gap, re-selects the gap, and adjusts the driving strategy; Update the candidate trajectory set, re-sample the candidate trajectories according to the new target gap, and repeat the planning process (sampling - evaluation - selection).
[0042] As described above, similar technical solutions can be derived from the solution content given in combination with the drawings and descriptions. Any solution content that does not depart from the structure of the present invention still falls within the scope of the technical solutions of this application.
Claims
1. An interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning, characterized in that: The decision-making is divided into two modes: Stackelberg game and Bayesian Stackelberg game based on the integrity of information. If all relevant vehicles provide interaction intentions through V2V, the Stackelberg game is adopted. If the intentions of some vehicles are opaque, the opponent's behavior needs to be inferred through sensors and probability models to trigger the Bayesian Stackelberg game; Specifically, it includes the following steps: Step (1) Environment perception and data collection: Perform V2V communication detection, receive the BSM of surrounding vehicles in real time, extract data, and obtain the states of surrounding vehicles through multi-sensor fusion; Step (2) V2V availability determination; Step (3) Select the game mode; If V2V is available, that is, a complete information scenario, select the Stackelberg game and execute the following step (3.1). If V2V is not available, that is, an incomplete information scenario, then select the Bayesian Stackelberg game and execute the following step (3.2); Step (3.1) Execute the selected Stackelberg game; The vehicle attempting to change lanes is the leader, and the obstacle vehicle is the follower. The decision-making process is carried out in stages. The leader acts first, and the follower then optimizes; Step (3.2) Execute the selected Bayesian Stackelberg game; The autonomous vehicle acts first as the leader, and the surrounding vehicles act as followers and respond according to the leader's actions, considering the types of vehicles at the same time; Step (4) Deep inverse reinforcement learning planning; Learn the reward function of human drivers from natural driving data and use the reward function to evaluate the humanity of candidate trajectories; Step (5) Planning; The dual-mode game decision module selects the target gap. The deep inverse reinforcement learning planning module samples candidate trajectories within the target gap range. For each candidate trajectory, use FIRL to predict the response trajectory of the FVTG. The reward network calculates the joint reward of the candidate trajectory and the FVTG response trajectory, and selects the trajectory with the highest reward as the planning result. If the planned joint trajectory causes a conflict, trigger the decision module to re-select the gap.
2. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 1, characterized in that: The above step (2) includes: Step (2.1) Communication module hardware status detection; Including signal strength detection and module activity check; Step (2.2) Communication quality index calculation: including packet loss rate and end-to-end delay; Step (2.3) Data consistency verification: including spatio-temporal alignment, error metric, and dynamic rationality verification; Step (2.4) V2V penetration rate evaluation; Step (2.5) Comprehensive availability determination: Weight and fuse all the above indicators, calculate and generate an availability score.
3. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 1, characterized in that: The above step (3.1) includes the following steps: Step (3.1.1) Participant division; The host vehicle is the vehicle attempting to change lanes. The host vehicle is the leader, and the obstacle vehicle is the vehicle behind in the target lane. The obstacle vehicle is the follower; Step (3.1.2) Utility function design; Consider the combined cost function of safety, traffic efficiency, and riding comfort to determine the Nash equilibrium; In the lane-changing behavior, the safety cost function of vehicle driving is manifested both horizontally and vertically; the formula of the safety cost function is as follows: (21) Among them, and respectively represent the safety cost functions in the horizontal and vertical directions; ; Ride comfort is related to lateral acceleration and longitudinal acceleration, and the expression of the comfort cost function is as follows: (22) In the formula, and are the weight coefficients of the lateral and longitudinal accelerations respectively, and are the lateral and longitudinal accelerations respectively; The expression of the cost function of efficiency is as follows: (23) In the formula, represents the maximum vehicle speed of the lane, represents the longitudinal speed of the host vehicle, represents the longitudinal speed of the vehicle in front in the current lane, represents the relative distance between the host vehicle and the vehicle behind in the target lane, represents the minimum safety distance; The cost function of the host vehicle is expressed as: (24) In the formula, is the weight coefficient of safety cost, is the weight coefficient of comfort cost, is the weight coefficient of efficiency cost; Step (3.1.3) Lane-changing decision-making; In the Stackelberg game, vehicles follow the principle of minimizing cost when making lane-changing decisions. Followers react to the leader's behavior based on their own cost functions and the influence of the leader's actions; The two-vehicle game optimization problem is expressed as: (25) (26) (27) (28) In the formula, represents the possible acceleration, represents the optimal acceleration of the host vehicle, represents if vehicle 1 is changing lanes, represents if changing lanes is optimal for vehicle 1, , represents the behavior decision of the vehicle, , represents the total cost function of the vehicle, represents the optimal decision of vehicle 2 under the influence of vehicle 1, , is the set of possible actions of the vehicle; Step (3.1.4) Output the target gap to the deep inverse reinforcement learning planning module.
4. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 1, characterized in that: The said step (3.2) Execute the selective Bayesian Stackelberg game; The autonomous vehicle acts first as the leader, and the surrounding vehicles act as followers and respond according to the leader's actions, while considering the types of vehicles; Specifically, it includes the following steps: Step (3.2.1) Define the game roles and order; The host vehicle acts as the leader and acts first. The vehicle behind in the target lane acts as the follower and selects a response strategy according to the leader's actions; The host vehicle selects a strategy based on the probability distribution of the follower types ; After observing the actions of the host vehicle, the follower selects an optimal response according to its own type ; Step (3.2.2) Model incomplete information; It includes a type space, a probability distribution, and a utility function; the host vehicle knows its own type but only has a prior distribution of the type of the FVTG , and the FVTG knows its own type , but is unsure of the specific strategy of the leader; The utility of the host vehicle needs to consider the response strategy of the followers and its type distribution: ; FVTG selects the optimal response according to the actions of the EVs and its own type: ; Step (3.2.3) Utility function design; Design the utility function by integrating safety, efficiency, and traffic friendliness; Step (3.2.4) Solve the Bayesian Stackelberg equilibrium; Step (3.2.5) Probability update and dynamic adjustment; The host vehicle predicts the real-time behavior of the follower through GMM, and updates the distribution of the follower type using Bayes' theorem, as shown in the following formula: (35) Among them, Generated by the SIDM model; R-R diagram combining time headway THW and acceleration, to update the aggressiveness of the follower in real time , for adjusting type probabilities; Step (3.2.6) Output the target gap to the deep inverse reinforcement learning planning module.
5. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 4, characterized in that: The said step (3.2.3) includes: Step (3.2.3.1) Safety cost function design; During the lane change process, for vehicle i, its driving safety within a certain prediction range h is defined by comparing the distance between vehicles and a predefined threshold safety distance , as well as the lane positioning of each vehicle and ; If the relative lane difference and are for all , then vehicle i is safe; if a vehicle is crossing a lane boundary, the vehicle may have multiple lanes; in this case, the safety of vehicle i can be ensured only when all relevant vehicles meet the definition of driving safety; In the above definition, is the relative lane difference ; is the distance between vehicle i and vehicle j at time t, as follows: (29) where \(h>0\) is the time interval, and are the longitudinal driving speeds of vehicle \(i\) and vehicle \(j\) at time \(t\); Many alternative safety metrics have been used to quantify the safety level, such as time headway (THW), time to collision (TTC), and the distance between vehicles; in addition, when a dangerous situation occurs, the driver's deceleration reaction time also plays a key role; In summary, these factors are integrated into the required deceleration, as shown in the following formula: (30) That is, under certain conditions, the deceleration required for the following vehicle to avoid a collision , represents the distance between vehicles, and represent the speeds of the front and rear vehicles, represents the driver's reaction time, represents the acceleration of the following vehicle, represents the maximum braking deceleration; the driving safety index is determined as: (31) In the formula, is the weighting factor; Step (3.2.3.2) Traffic friendliness cost function design; Traffic friendliness is represented by restricting the frequency of lane changes; lane changes are only initiated when the current lane cannot meet the driving function; if the required deceleration in the future is less than the threshold of the current lane , then a lane change penalty will be effectively imposed, and the traffic friendliness factor can be expressed as: (32) In the formula, is the weight factor, is the threshold deceleration to ensure driving safety, is the predicted required deceleration calculated according to the future state of the vehicle, is the sign function; Step (3.2.3.3) Efficiency cost function design; The driving efficiency of vehicle i is determined by the difference between the current speed and the future speed . The speed for the long-term stable following state, the expression of the efficiency cost function is: For the vehicle speed in the long-term stable car-following state, the expression of the efficiency cost function is: (33) In the formula, represents the maximum speed that vehicle i can freely accelerate to when it is greater than the safe following distance, and it is a weight factor.
6. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 4, characterized in that: The said step (3.2.4) includes: Step (3.2.4.1) Backward induction; Follower response modeling, for each possible leader strategy and follower type , solve for the follower's optimal response ; Leader strategy optimization, with the host vehicle as the leader, it is necessary to maximize its own expected utility, considering the uncertainty and response of the follower type, as follows: (34) In the formula, is the leader's strategy, is the follower type, is the prior distribution of the follower, is the optimal response of the follower, is the utility of the host vehicle; Step (3.2.4.2) Bilevel optimization problem; Upper-level problem, leader's optimization strategy ; Lower-level problem, for each , the follower solves ; Use SQL to solve the bilevel optimization problem.
7. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 1, characterized in that: The said step (4) includes: Step (4.1) Data preparation and preprocessing; Use the NGSIM public driving dataset to screen out expert trajectory segments containing highway lane-changing scenarios; construct a dataset suitable for DIRL training to ensure that the input format is unified and the features are complete; Step (4.2) Adopt sampling-based deep inverse reinforcement learning; According to the sampling-based DIRL to solve the problem of the solution difficulty of DIRL, according to the principle of maximum entropy, the probability of selecting a trajectory is proportional to the natural exponent of its reward; Step (4.3) Deep inverse reinforcement learning planning; Through sampling, evaluation, and selection, select the trajectory with the highest reward value.
8. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 7, characterized in that: The said step (4.2) solves the problem of the difficulty in solving DIRL according to the principle of maximum entropy, and the probability of selecting a trajectory is proportional to the natural exponent of its reward, as shown in the following formula: (36) wherein, represents the probability that the trajectory τ is selected, is the reward of the trajectory τ, is the parameter of the reward function, is called the normalization function or partition function, and D is the set of all possible trajectories that the agent can choose; The continuous and huge state space makes it difficult to calculate the partition function in the lane-changing scenario. The following partition function can be adopted to approximately represent it by sampling candidate trajectories in the state space: (37) Among them, represents all sampled candidate trajectories, is the trajectory 's reward; then the probability that the trajectory τ is selected is: (38) Among them, all the trajectories in have the same initial state as the trajectory τ; Using maximum entropy inverse reinforcement learning to maximize the likelihood of the expert demonstration trajectories by adjusting the parameters of the reward function as follows: as follows: (39) where, E represents all expert demonstration trajectories; then the objective function of IRL is given by the following formula: (40) Among them, E is all expert demonstration trajectories, is a single expert demonstration trajectory, is the trajectory reward, ; is the locus probability of being selected; The gradient of the parameters of the derived reward function is: (41) In the formula, is the trajectory and the selected probability.
9. The interactive perception autonomous driving method based on dual-mode game and deep inverse reinforcement learning according to claim 7, characterized in that: The said step (4.3) includes: Step (4.3.1) Candidate trajectory sampling; It includes a longitudinal adjustment stage and a lateral lane-changing stage; The longitudinal adjustment phase samples high-dimensional parameters, the target longitudinal position within the time interval , speed , and adjustment time ; and uses the following fifth-degree polynomial to fit the longitudinal position of the trolley in the first stage under the S-L coordinate system: (42) Among them, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, and t is the time variable; Then the speed and acceleration of the host vehicle are: (43) (44) The boundary conditions include the initial state and the target state , then the coefficients of the polynomial can be expressed as: (45) wherein, is the initial moment, is the target moment, is the initial longitudinal position, is the initial speed, is the initial acceleration, is the target longitudinal position, is the target speed, is the target acceleration; In the lateral lane-changing stage, the lateral trajectory of the lane change is fitted by formula (45); the L coordinate in the S-L coordinate system is expressed as: (46) Among them, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial, is the coefficient of the fifth-degree polynomial; Candidate trajectories can be generated by sampling high-level driving intentions with a sampling range for the required time being , a sampling interval of 1 s, and a sampling range for the target speed being , a sampling interval of 0.5 m / s, where is the speed of the vehicle behind the ego vehicle in the target lane and the vehicle closest to the ego vehicle among the vehicles behind, and the sampling range for the target longitudinal position is m, a sampling interval of 2 m, where is the position of the vehicle behind the ego vehicle in the target lane at time and is the position of the vehicle in front of the ego vehicle in the target lane at time ; any candidate trajectory that results in a collision will be removed; Step (4.3.2) Interactive perception prediction; In the lane-changing scenario, the host vehicle mainly affects the vehicle behind the target gap. The vehicle behind the target gap is represented by FVTG. The trajectory of FVTG is predicted using feature-based inverse reinforcement learning; FIRL learns the internal reward function of FVTG from driving data and finds the trajectory that maximizes the reward function as the predicted trajectory; the lane-changing scenario is regarded as a car-following scenario where the leading vehicle in FVTG suddenly changes; FVTG can recognize the lane-changing intention of the ego vehicle and regard it as the new leading vehicle when the ego vehicle signals a lane change using its turn signal or historical trajectory; In this scenario, the feature selection of FVTG is as follows: 1) Acceleration: (47) where, acc is the instantaneous acceleration of FVTG; 2) Speed relative to the ego vehicle: (48) Among them, is the relative speed between the FVTG and the host vehicle; 3) Relative distance to the ego vehicle: (49) where d is the actual relative distance between the FVTG and the host vehicle, represents the expected relative distance to the ego vehicle; the distance between the FVTG and the LVTG is regarded as the desired relative distance; 4) Collision penalty: (50) In feature-based inverse reinforcement learning, the reward function is as follows: (51) wherein is the weight parameter of the acceleration feature, is the weight parameter of the relative velocity feature, is the weight parameter of the relative distance feature, is the weight parameter of the collision penalty feature; Given the reward function, the optimal trajectory of FVTG is planned using MPC; that is, the reward function of the traffic vehicle FVTG is learned through FIRL, and its trajectory is generated through model predictive control MPC; Step (4.3.3) Reward network evaluation; A reward network is used to represent the reward function used by humans for planning, and sampling-based DIRL is used to train the reward network; Step (4.3.4) Sort the candidate trajectories according to the reward value, and select the trajectory with the highest reward value as the final planning result.
10. The interactive perception autonomous driving method based on bimodal game and deep inverse reinforcement learning according to claim 1, characterized in that: The said step (5) includes: Step (5.1) Conflict detection; Step (5.2) Reward evaluation and threshold setting; The reward network will output a comprehensive score for each candidate trajectory, including safety, efficiency, comfort, and interaction cost; safety threshold, if the score is lower than the safety threshold, it is determined as an infeasible trajectory; dynamic adjustment, the threshold can be dynamically adjusted according to the scenario, for example, increasing the safety requirements in dense traffic; Step (5.3) Logic to trigger re-decision; When a conflict or insufficient reward score is detected, the following process is triggered for adjustment: Conflict marking, the planning module marks the conflict trajectory as "high risk" and records the conflict type; The decision-making module receives the conflict information, re-evaluates the feasibility of the current target gap, re-selects the gap, and adjusts the driving strategy; Update the candidate trajectory set, re-sample candidate trajectories according to the new target gap, and repeat the planning process.
Citation Information
Patent Citations
Freight vehicle interactive game lane change decision-making method based on lane change decision-making system
CN114882705A
Automatic driving vehicle lane changing behavior vehicle road collaborative decision algorithm based on Bayesian game
CN115056798A
Blind zone intersection support method and system using V2V communication
CN115148048A
Automatic driving vehicle lane changing trajectory planning method based on multi-person game
CN116185027A
Personalized driver-based anthropomorphic lane changing trajectory optimization method
CN116534055A
Cited By
Robot autonomous navigation method based on reinforcement learning in complex dynamic environment
CN120558244A
Vehicle trajectory prediction method and device based on space-time cooperation and game driving
CN121019599A
Self-driving vehicle hierarchical game decision-making method for low-speed afflux scene
CN121165751A
Multi-driving-style high-risk automatic driving cut-in scene test method and system
CN121387751A
Hybrid traffic interaction behavior trajectory prediction method and device, and storage medium
CN121457757A