A method and device for establishing a high-speed entrance ramp forced merging decision model based on game theory

By adopting a game theory-based two-layer decision model in the entrance ramp scenario, considering the comprehensive behavior of main lane vehicles, the problem that the lane change behavior of main lane vehicles in the existing technology is not fully considered, and the accuracy of the model and the decision-making ability of the autonomous driving system are improved.

CN119229675BActive Publication Date: 2025-05-09ZHEJIANG UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411218461.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-05-09
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

In the prior art forced merge decision model in the inlet ramp scenario, the lane change behavior of the main lane vehicles is not fully considered, resulting in insufficient model prediction success rate, and data imbalance problem affects modeling accuracy.

Method used

A two-layer decision model based on game theory is adopted, and a more comprehensive decision model for main lane vehicles is established through data processing and model calibration. Strategies such as acceleration, keeping the same, and polite avoidance of main lane vehicles are considered. In polite avoidance, it is divided into direct avoidance of slowing and direct avoidance of lane change. Further decisions are made as to slowing down and direct avoidance of lane change.

Benefits of technology

It improves the accuracy of the vehicle simulation system in the entrance ramp scenario, enhances the credibility of control algorithm detection, can more accurately predict the behavior of main road vehicles and ramp vehicles, and improves the decision-making ability of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229675B_ABST
    Figure CN119229675B_ABST
Patent Text Reader

Abstract

A method and device for establishing a forced merging decision model for a high-speed entrance ramp based on game theory, the method comprising the following steps: first, data processing is performed, and data enhancement is performed using synthetic minority class oversampling technology to solve the problem of model inaccuracy caused by data imbalance. Subsequently, a two-layer decision model is established based on game theory. When the main road vehicle strategy output by the upper-layer decision model is polite avoidance, the lower-layer game model is started to further decide whether to slow down and go straight to avoid or change lanes to avoid. Next, the model is calibrated to determine the optimal model parameters to minimize the difference between the actual merging decision in the data set and the merging decision predicted by the model. Finally, the confusion matrix is ​​used to verify the model accuracy and evaluate the model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving, and in particular to a method and device for establishing a high-speed entrance ramp forced merging decision model based on game theory. Background Art

[0002] In recent years, autonomous driving technology has developed rapidly. Autonomous vehicles (AVs) have gradually become an important part of traffic roads and will soon form mixed traffic situations. AVs help improve the efficiency and safety of the traffic system and alleviate traffic congestion. The application of complex traffic scenarios is the main challenge facing current autonomous driving technology, especially the high-speed entrance ramp merging scenario. The merging vehicles will affect the traffic vehicles on the main road, causing traffic congestion and even posing a safety threat. Therefore, autonomous driving technology for entrance ramp traffic is of great significance.

[0003] There are many vehicle control algorithms in the on-ramp scenario, such as rule-based control algorithms, model predictive control, game theory, and reinforcement learning. Currently, the accuracy of the algorithm is mainly verified through a simulation environment. Commonly used vehicle simulation software tools include CarSim, VISSIM, and IPG-CarMarker. However, there is a certain gap between the simulation environment and the real world. For example, when verifying the effectiveness of the algorithm for ramp vehicles, the model of the main road vehicle in the simulation environment is usually assumed to be a simple rule-based model (Intelligent Driver Model, IDM), which is often not the case in reality. How to ensure the accuracy of vehicle modeling in the simulation environment has become a challenge in the current field of autonomous driving.

[0004] In recent years, there are many models for forced merging of ramp vehicles, such as the gap selection of ramp vehicles, the decision whether to merge, when to merge, and the motion planning modeling of the decision to merge. However, there are few technologies for modeling vehicles on the main road, which is equally important. There are four main reactions of vehicles on the main road: acceleration, deceleration, keeping the acceleration constant, and changing lanes. Most of the current technologies only consider the acceleration and deceleration behavior of the vehicle, ignoring the behavior of keeping the acceleration constant and changing lanes to avoid, especially the lane changing behavior. Since the lane changing behavior of vehicles on the main road in the actual data set only accounts for about 3% of the entire data set, this extreme data imbalance is the main reason affecting the accuracy of modeling. In the current methods, it is difficult to find a set of model parameters that can achieve ideal results for both the prediction success rate of changing lanes and keeping straight. This is the difficulty and gap of the current ramp merging autonomous driving technology.

[0005] Many modeling methods do not consider the interaction between vehicles, only absolute safety, and the models are not accurate enough. Game theory is a very good choice for dealing with traffic interaction issues, and can effectively simulate the game between drivers. However, when using game theory methods to conduct the game between main road vehicles and ramp vehicles, when the main road vehicle chooses to change lanes, a new game will be generated. Currently, many methods only consider one game and ignore this new game.

[0006] In summary, the current technology of the forced merging decision model in the on-ramp scenario has the following limitations: 1) The main road behavior is not considered comprehensively, especially the lack of technology for lane changing behavior of vehicles on the main road; 2) The lane changing data of vehicles on the main road is unbalanced with other behavior data, and it is difficult to find a set of model parameters that can achieve reasonable results for both the prediction success rate of lane changing and staying straight; 3) Failure to consider the lane changing behavior of vehicles on the main road will generate a new game. Summary of the invention

[0007] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provides a method for establishing a high-speed entrance ramp forced merging decision model based on game theory.

[0008] The present invention models the decision-making of ramp vehicles and main road vehicles in the scenario of forced merging of the entrance ramp. Through game theory modeling, it helps people better understand the behavior of main road drivers and ramp drivers in the scenario of entrance ramp merging. Secondly, the model is applied to the vehicle simulation system to further improve the accuracy of the vehicle simulation system and increase the credibility of the control algorithm detection. Finally, it can be integrated into the prediction module of the main road and ramp autonomous driving vehicles to predict the behavior of main road vehicles and ramp vehicles.

[0009] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0010] First, data processing is performed, and synthetic minority class oversampling technology is used for data enhancement. Secondly, a two-layer decision model is established. The upper layer is a game decision model between main road vehicles and ramp vehicles. The main road vehicle has three strategies: acceleration, unchanged, and polite avoidance. Among them, the polite avoidance strategy includes two: slow down and go straight to avoid and change lanes to avoid. No distinction is made in the upper layer decision. The lower layer is a game decision model between main road vehicles and the rear vehicles in the adjacent road. The main road vehicle has two strategies: slow down and go straight to avoid and change lanes to avoid. When the upper layer decision model outputs the strategy of polite avoidance, the lower layer game model is started to further decide whether to slow down and go straight to avoid or change lanes to avoid. Next, the model is calibrated to determine the optimal model parameters to minimize the difference between the actual merged decision in the data set and the merged decision predicted by the model. Finally, the confusion matrix is ​​used to verify the model accuracy and evaluate the model performance.

[0011] A method for establishing a high-speed entrance ramp forced merging decision model based on game theory of the present invention specifically comprises the following steps:

[0012] The first step is data processing. First, the data is classified and extracted according to the different behaviors of the vehicles. Then the SMOTE method is used for data enhancement. 70% of the data is used for model calibration and 30% for model verification.

[0013] The second step is to establish a decision-making model, design a profit function, and establish a two-level decision-making model based on game theory;

[0014] The third step is model calibration, which uses the gradient descent algorithm to determine the model parameters to minimize the difference between the decision results of the data set and the decision results predicted by the model;

[0015] In the fourth step, model validation, the confusion matrix was used to evaluate the performance of the model based on the parameter estimates obtained during the calibration process in the third step.

[0016] The first step of data processing described in the technical solution is to first classify and extract data according to the different behaviors of the vehicle, and then use the SMOTE method to enhance the data. 70% of the data is used for model calibration and 30% is used for model verification.

[0017] Since the two-layer decision model proposed in the present invention has two games, the data is processed twice. The result of the first processing is used for the game between the main road vehicle and the ramp vehicle, and the result of the second processing is used for the game between the main road vehicle and the rear vehicle of the adjacent road.

[0018] The first data processing extracts the data of vehicles on the main road and ramps. The trajectory data of ramp vehicles is divided into merging segments and waiting segments, and the data of vehicles on the main road is divided into acceleration segments, unchanged segments, and polite avoidance segments. The second data processing extracts the data of vehicles on the main road and adjacent roads. The vehicles on the main road are divided into deceleration straight-ahead segments and lane-changing segments, and the data of vehicles on adjacent roads are divided into avoidance segments and non-avoidance segments.

[0019] After data extraction and classification, the imbalance between different categories of data will affect the final accuracy of the model. The present invention uses the SMOTE method to enhance the data to balance the data. The SMOTE method expands the minority class data by synthesizing minority class samples. In the feature space, this method generates new synthetic samples by interpolating minority class samples and their neighboring samples, thereby balancing the category distribution in the data set.

[0020] First, identify the minority samples in the dataset and separate them from the dataset. For an unbalanced dataset, the number of minority samples is much lower than that of majority samples. Use the K-Nearest Neighbors (KNN) algorithm to calculate the K nearest neighbor samples of each minority sample in the feature space. For each minority sample A, randomly select a neighbor sample (such as B1) from its K closest minority samples B1, B2, ..., BK, and generate a new synthetic sample between sample A and the neighbor sample B1. The synthetic sample is generated by linearly interpolating the feature difference between A and B1:

[0021] S=A+σ×(B1-A) (1)

[0022] Where S represents the synthetic sample, and σ is a random number between 0 and 1, which is used to control the position of the interpolation.

[0023] This process is repeated until a sufficient number of synthetic samples are generated, and the generated synthetic samples are added to the original dataset to form a new balanced dataset so that the number of minority class samples reaches the desired balance target.

[0024] The second step in the technical solution is to establish a decision-making model, design a profit function, and establish a two-layer decision-making model based on game theory.

[0025] The upper-level game theory model is specifically designed as follows:

[0026] 1) Determine the game participants: the merging vehicle (MV) and the following vehicle (FV);

[0027] 2) Determine the strategy set of game participants: S MV {Waiting for merging, merging}, S FV {speed up, maintain stability, and give way politely};

[0028] 3) Revenue function design: The design goal of MV is to complete the merger as quickly as possible while ensuring safety, and the design goal of FV is to ensure safety and minimize speed fluctuations.

[0029] The FV profit function is designed as follows:

[0030] When the MV chooses to merge and the FV chooses to yield politely, the predicted acceleration required by the FV is calculated as follows:

[0031]

[0032] Where w is the width of MV, d1 and d2 are the distances from FV to the left and right rear of MV, respectively, and s eis an intermediate variable, s* represents the ideal distance between FV and the preceding vehicle, s0 represents the minimum distance between FV and the preceding vehicle, T represents the ideal headway, and v F 、v M Respectively represent the speed of FV and MV at the current moment, a max represents the maximum acceleration of FV, d comfort represents the comfortable deceleration of FV, v0 represents the ideal speed of FV, Acc FV_md This is the predicted acceleration required for FV.

[0033] At this time, the profit function of FV is:

[0034] U FV_md =α1+λ1Acc FV_md (5)

[0035] Where α1 and λ1 are the parameters to be calibrated.

[0036] When MV chooses to merge and FV chooses to accelerate, the predicted acceleration required by FV is calculated as follows:

[0037] The speed of the MV at the predicted merging time can be calculated based on the speed and acceleration of the MV at the current moment and the remaining distance in the acceleration lane:

[0038]

[0039] In the formula, a M represents the acceleration of MV at the current decision moment, and RD represents the remaining distance of MV in the acceleration lane.

[0040] From this, the remaining time of MV in the acceleration channel can be calculated:

[0041]

[0042] The predicted acceleration required for FV is calculated as follows:

[0043]

[0044] Where v' F represents the speed of predicting FV, a F represents the acceleration of FV at the current decision moment, X and X' represent the distance between FV and MV at the current decision moment and the prediction moment, respectively, t b is the reaction time, generally 2s.

[0045] At this time, the profit function of FV is:

[0046] U FV_ma =α2+λ2Acc FV_ma (11)

[0047] Where α2 and λ2 are the parameters to be calibrated.

[0048] When the MV selections are merged and the FV selections remain unchanged, the predicted acceleration required by the FV is calculated as follows:

[0049]

[0050] As shown in formula (12), if the current MV speed is greater than the FV speed, the MV will complete the merger without interfering with the FV. If the MV speed is lower than the FV speed, the MV will force the FV to slow down.

[0051] At this time, the profit function of FV is:

[0052] U FV_mdn =α3+λ3Acc FV_mdn (13)

[0053] Where α3 and λ3 are the parameters to be calibrated.

[0054] When MV chooses to wait, the calculation of FV's payoff function is similar.

[0055] The profit function of MV is designed as follows:

[0056] When the MV chooses to merge and the FV chooses to give way politely, the MV can merge at a comfortable acceleration. The predicted acceleration required by the MV is:

[0057] Acc MV_md =Acc comfort (14)

[0058] Where Acc comfort It is the comfortable merging acceleration of MV.

[0059] The profit function of MV is:

[0060] U MV_md =β1+η1Acc MV_md (15)

[0061] Where β1 and η1 are the parameters to be calibrated.

[0062] When the MV chooses to merge and the FV chooses to accelerate or keep the same strategy, the MV needs to travel at the maximum acceleration to ensure that it reaches the merging point before the FV. The acceleration required by the MV is as follows:

[0063] Acc MV_ma =Acc MV_mdn =Acc max (16)

[0064] Where Accmax Indicates the maximum acceleration of the MV.

[0065] At this time, the profit functions corresponding to MV are:

[0066] U MV_ma =β2+η2Acc MV_ma (17)

[0067] U MV_mdn =β3+η3Acc MV_mdn (18)

[0068] Where β2, β3, η2, η3 are the parameters to be calibrated.

[0069] When the MV chooses to wait and the FV chooses to give way politely, the MV can merge with the FV at a comfortable acceleration after determining the FV avoidance strategy. The predicted acceleration required by the MV is:

[0070] Acc MV_wd =Acc comfort (19)

[0071] At this time, the profit function of MV is:

[0072] U MV_wd =β4+η4Acc MV_wd (20)

[0073] Where β4 and η4 are the parameters to be calibrated.

[0074] When MV chooses to wait and FV chooses to accelerate or maintain the status quo, MV needs to wait for FV to overtake before merging. The waiting time of MV is calculated as follows:

[0075]

[0076] At this time, the predicted acceleration required by MV is calculated as follows:

[0077]

[0078] The profit function of MV is:

[0079] U MV_wa =β5+η5Acc MV_wa (twenty three)

[0080] U MV_wdn =β6+η6Acc MV_wdn (twenty four)

[0081] Where β5, β6, η5, η6 are the parameters to be calibrated.

[0082] In summary, the design of the upper-level game theory payoff function is completed.

[0083] The lower-level game model is specifically designed as follows:

[0084] 1) Determine the game participants: the main lane vehicle (FV) and the adjacent lane rear vehicle (Lag);

[0085] 2) Determine the strategy set of game participants: S FV {Change lane, go straight}, S Lag {slow down to avoid, do not avoid};

[0086] 3) Benefit function design: The design objectives of FV and Lag are to ensure safety and minimize speed fluctuations.

[0087] The profit function design of FV and Lag is as follows:

[0088] The deceleration of the vehicle to avoid a collision is:

[0089]

[0090] In the above formula, v B represents the speed of the vehicle ahead, v A Indicates the speed of the rear vehicle, s B represents the position of the vehicle ahead, s A Indicates the position of the rear vehicle, l A Indicates the length of the vehicle behind;

[0091] The profit function of FV and Lag is calculated as follows:

[0092]

[0093] In the formula, i represents the decision of FV, j represents the decision of Lag, and a FV,ij represents the acceleration of FV when FV makes decision i and Lag makes decision j, a Lag,ij represents the acceleration of Lag when FV makes decision i and Lag makes decision j. μ1, μ2, μ3, μ4, ψ1, ψ2 and ψ3 are the parameters to be optimized.

[0094] In summary, the design of the lower-level game theory payoff function is completed.

[0095] The two game theory models constitute a two-layer decision-making model. When MV enters the acceleration channel, the upper-layer game theory matrix begins to be constructed. When FV chooses to give way politely after bargaining with MV, FV then bargains with Lag to decide whether to change lanes or slow down and go straight.

[0096] The third step described in the technical solution, model calibration, uses a gradient descent algorithm to determine the model parameters to minimize the difference between the decision results of the data set and the decision results predicted by the model.

[0097] The objective function of the upper-level game theory model calibration is as follows:

[0098]

[0099] In the formula, F i and They represent the decision results of the actual data set FV and the decision results predicted by the model, respectively, and M i and They respectively represent the decision results of the MV actual data set and the decision results predicted by the model, i represents the index of the event in the data set, and n represents that a total of n events are extracted from the data set.

[0100] The objective function for calibrating the underlying game theory model is as follows:

[0101]

[0102] Where, L i and They represent the decision results of the Lag actual data set and the decision results predicted by the model respectively, and m means that a total of m events are extracted from the data set.

[0103] In order to minimize the objective function, the present invention adopts the gradient descent algorithm. The gradient descent algorithm is a commonly used iterative search algorithm that gradually approaches the optimal solution by moving along the negative direction of the function gradient of the current point. This method is particularly suitable for processing models involving a large number of parameter estimates and high complexity due to its high computational efficiency and relatively simple implementation. In the present invention, due to the complexity of the model and the requirement for the number of parameters, the gradient descent method is selected as the preferred method for calibrating the model. This not only improves the efficiency of the calculation, but also ensures the accuracy and robustness of the model.

[0104] The parameter calibration of the double-layer decision model in the present invention adopts a separate calibration method, and the two processing results of the data set are calibrated to two game theory models respectively, and finally two sets of calibration parameters are obtained. The two game theory models are combined to obtain the double-layer decision model.

[0105] The fourth step described in the technical solution, model validation, uses a confusion matrix to evaluate the performance of the model based on the parameter estimates obtained during the third step calibration process.

[0106] In order to verify the effectiveness of the proposed two-level decision model, this paper uses a series of performance indicators to evaluate the model. These performance indicators are mainly based on the confusion matrix analysis method, which can deeply reveal the model's ability to predict behavior. The confusion matrix contains many important performance indicators, including:

[0107] 1) True Positive (TP): The decision predicted by the model is consistent with the decision of the actual data set, reflecting the correct prediction ability of the model;

[0108] 2) False Positive (FP): The decision result predicted by the model is inconsistent with the decision result of the actual data set, indicating that the model may have misjudgment;

[0109] 3) Detection Rate: The proportion of correctly predicted events in all actual events, which measures the sensitivity of the model;

[0110] 4) False Alarm Rate: The ratio of incorrectly predicted events to all events in the data set, indicating the false alarm frequency of the model.

[0111] In addition, in order to more comprehensively understand the performance of the proposed model, the present invention independently evaluates the model performance of each strategy to further verify the accuracy of the model.

[0112] A second aspect of the present invention relates to a device for establishing a highway entrance ramp forced merging decision model based on game theory, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a method for establishing a highway entrance ramp forced merging decision model based on game theory of the present invention.

[0113] A third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements a method for establishing a freeway entrance ramp forced merging decision model based on game theory of the present invention.

[0114] The beneficial effects of the present invention are mainly manifested in:

[0115] (1) The present invention uses synthetic minority oversampling technology (SMOTE) to enhance the data, which effectively solves the problem of insufficient lane-changing behavior data of vehicles on the main road, so that the model can still maintain good prediction performance when facing unbalanced data, thereby improving the accuracy of the model;

[0116] (2) The present invention establishes a two-layer decision model, which not only considers the acceleration, maintenance and polite avoidance strategies of the main road vehicles when the ramp vehicles merge, but also further subdivides the polite avoidance into deceleration and straight-ahead avoidance and lane change avoidance, which fully covers the possible behaviors of the main road vehicles. In addition, the new game is considered when the main road vehicles change lanes, which is more in line with the actual driving scenario.

[0117] (3) The two-layer decision model of the present invention not only helps people understand driving behavior in ramp merging scenarios, but can also be integrated into the prediction modules of main road and ramp autonomous driving vehicles to predict the behaviors of main road vehicles and ramp vehicles in real time, providing more intelligent decision support for the autonomous driving system. Finally, it can also be embedded in the vehicle simulation system, increasing the credibility of the vehicle simulation system in complex traffic scenarios, and providing reliable support for the subsequent development and verification of autonomous driving algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0118] Figure 1 The present invention shows the process of establishing a method for a compulsory merging decision model for a high-speed entrance ramp based on game theory.

[0119] Figure 2 A parameter calibration flow chart of the upper-level game theory model of the present invention is shown.

[0120] Figure 3 A parameter calibration flow chart of the underlying game theory model of the present invention is shown. DETAILED DESCRIPTION

[0121] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings.

[0122] Example 1

[0123] Reference Figure 1 to Figure 3 , a high-speed entrance ramp forced merging decision model based on game theory, firstly extracts and classifies data, then uses synthetic minority class oversampling technology for data enhancement, and then establishes a two-layer decision model, the upper layer is the game decision model between main road vehicles and ramp vehicles, and the lower layer is the game decision model between main road vehicles and adjacent lane vehicles. Then the model is calibrated and the model parameters are determined to minimize the difference between the merging decision of the data set and the merging decision predicted by the model, and finally the model is validated to evaluate the model performance.

[0124] Reference Figure 1 , showing the overall workflow of the present invention.

[0125] A method for establishing a high-speed entrance ramp forced merging decision model based on game theory includes the following steps:

[0126] S1. Classify and extract data according to different behaviors of vehicles, and use SMOTE method to enhance data.

[0127] S2. Design the profit function and establish a two-level decision-making model based on game theory.

[0128] S3. Use the gradient descent algorithm to determine the model parameters to minimize the difference between the decision results of the dataset and the decision results predicted by the model.

[0129] S4. Use confusion matrix to evaluate the performance of the model.

[0130] In S1 described in the figure, data is first classified and extracted according to different behaviors of vehicles, and then the SMOTE method is used for data enhancement. 70% of the data is used for model calibration and 30% for model verification. Since there are two games in the two-layer decision model proposed in the present invention, the data is processed twice. The result of the first processing is used for the game between the main road vehicle and the ramp vehicle, and the result of the second processing is used for the game between the main road vehicle and the rear vehicle of the adjacent road. The second data processing implementation is taken as an example for explanation.

[0131] The present invention mainly uses the python pandas library for data processing. First, the trajectory data of FV is extracted, and the speed time series data is segmented and linearized using the bottom-up algorithm. The slope of each segment is judged. If the slope is greater than 0.05g (g is the acceleration of gravity), the FV strategy corresponding to this segment of data is considered to be acceleration. If it is less than -0.05g, the FV strategy corresponding to this segment of data is considered to be deceleration. If the slope is between -0.05g and 0.05g, the FV strategy corresponding to this segment of data is considered to remain unchanged, and the lane change strategy is directly judged according to the position trajectory. Then the lane change data segment and the deceleration straight segment of FV are extracted, and the data of Lag vehicles are extracted and classified in the same way, and the above data is used as the data set for the calibration of the parameters of the second game theory model. . The acceleration data segment, the unchanged segment and the polite avoidance segment (the set of the deceleration avoidance segment and the lane change segment) of FV are extracted, and the data of MV are extracted and classified, and the above data is used as the data set for the calibration of the parameters of the first game theory model.

[0132] After completing the policy judgment on the event data segment, it will be marked with the corresponding policy label. Subsequently, these labeled event data segments will be stored in the form of a dictionary structure and saved in a JSON format file.

[0133] After data extraction and classification, the imbalance between different categories of data will affect the final accuracy of the model. The present invention uses the SMOTE method to enhance the data to balance the data. The SMOTE method expands the minority class data by synthesizing minority class samples. In the feature space, this method generates new synthetic samples by interpolating minority class samples and their neighboring samples, thereby balancing the category distribution in the data set.

[0134] S2 in the figure establishes a decision model, designs a payoff function, and uses game theory to establish a two-layer decision model. When MV enters the acceleration channel, the upper-layer game theory payoff matrix begins to be constructed. At this time, MV and FV play a game. The model calculates the payoff matrix based on the current state of the vehicles participating in the game, obtains the Nash equilibrium solution, and obtains the current optimal strategy solution. When FV's decision solution is to accelerate or remain unchanged, the strategy result is directly output. When FV's decision solution is to politely avoid, the second-layer game starts. At this time, the game participants become FV and Lag. The second game theory model obtains the optimal decision result for both parties based on the current state of FV and Lag's vehicles. FV finally decides whether to change lanes to avoid or slow down and go straight to avoid.

[0135] The upper-level game theory model is specifically designed as follows:

[0136] 1) Determine the game participants: the merging vehicle (MV) and the following vehicle (FV);

[0137] 2) Determine the strategy set of game participants: S MV {Waiting for merging, merging}, S FV {speed up, maintain stability, and give way politely};

[0138] 3) Revenue function design: The design goal of MV is to complete the merger as quickly as possible while ensuring safety, and the design goal of FV is to ensure safety and minimize speed fluctuations;

[0139] The lower-level game model is specifically designed as follows:

[0140] 1) Determine the game participants: the main lane vehicle (FV) and the adjacent lane rear vehicle (Lag);

[0141] 2) Determine the strategy set of game participants: S FV {Change lane, go straight}, S Lag {slow down to avoid, do not avoid};

[0142] 3) Benefit function design: The design objectives of FV and Lag are to ensure safety and minimize speed fluctuations;

[0143] S3, model calibration, described in the figure, uses a gradient descent algorithm to determine the model parameters to minimize the difference between the decision results of the data set and the decision results predicted by the model.

[0144] The objective function of the upper-level game theory model calibration is as follows:

[0145]

[0146] The objective function for calibrating the underlying game theory model is as follows:

[0147]

[0148] As described in Figure S4, model validation, the confusion matrix is ​​used to evaluate the performance of the model based on the parameter estimates obtained during the third step calibration process.

[0149] In order to verify the effectiveness of the proposed two-tier decision model, the present invention uses a series of performance indicators to evaluate the model. These performance indicators are mainly based on the confusion matrix analysis method, which can deeply reveal the ability of the model in predicting behavior. The confusion matrix contains many important performance indicators, including: true positive examples, false positive examples, detection rate, and false alarm rate. In addition, in order to more comprehensively understand the performance of the proposed model, the present invention independently evaluates the model performance of each strategy to further verify the accuracy of the model.

[0150] Reference Figure 2 , shows the parameter calibration flow chart of the upper-level game theory decision model of the present invention. In the figure, j represents the number of iterations. First, the initial parameter values ​​are set, including the number of iterations j = 0, and the initial values ​​of the calibration parameters α, β, λ, η. The profit matrix is ​​calculated based on the trajectory data in the data set, the Nash equilibrium solution is calculated, and the prediction strategy of the model is obtained. and Extracting the actual policy F from the NGSIM dataset i and M i , and calculate the error formula err, and use gradient descent to optimize parameters according to the current error value until the error is lower than the set threshold, then the cycle stops and the calibration is completed.

[0151] Reference Figure 3 , shows the parameter calibration flow chart of the lower-level game theory decision model of the present invention. In the figure, j represents the number of iterations. First, the initial parameter values ​​are set, including the number of iterations j = 0, and the initial values ​​of the calibration parameters μ and ψ. The profit function is calculated based on the trajectory data in the data set to obtain the model prediction strategy and Extracting the actual policy F from the NGSIM dataset i and L i , and calculate the error formula err, and use gradient descent to optimize parameters according to the current error value until the error is lower than the set threshold, then the cycle stops and the calibration is completed.

[0152] Example 2

[0153] The present embodiment relates to a device for establishing a highway entrance ramp forced merging decision model based on game theory, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement a method for establishing a highway entrance ramp forced merging decision model based on game theory of embodiment 1.

[0154] Example 3

[0155] This embodiment relates to a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a method for establishing a freeway entrance ramp forced merging decision model based on game theory in embodiment 1 is implemented.

[0156] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms described in the embodiments. The protection scope of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A method for establishing a high-speed entrance ramp forced merging decision model based on game theory, characterized in that: The synthetic minority over-sampling technique (SMOTE) is used for data enhancement, the two-layer decision model is established by game theory, and the model parameters are calibrated by the gradient descent optimization algorithm. Finally, the model is verified by the test data set. The following steps are included: The first step is data processing. First, the data is classified and extracted according to the different behaviors of the vehicles. Then the SMOTE method is used for data enhancement. 70% of the data is used for model calibration and 30% for model verification. The second step is to establish a decision-making model, design a profit function, and establish a two-level decision-making model based on game theory; The third step is model calibration, which uses the gradient descent algorithm to determine the model parameters to minimize the difference between the decision results of the dataset and the decision results predicted by the model; The fourth step, model validation, uses a confusion matrix to evaluate the performance of the model based on the parameter estimates obtained during the calibration process in the third step; In the first step, the data is processed twice. The result of the first processing is used for the game between the main road vehicle and the ramp vehicle, and the result of the second processing is used for the game between the main road vehicle and the rear vehicle of the adjacent road. The first data processing extracts the data of vehicles on the main road and ramps. The trajectory data of ramp vehicles is divided into merging segments and waiting segments. The data of vehicles on the main road is divided into acceleration segments, unchanged segments and polite avoidance segments. The second data processing extracts the data of vehicles on the main road and vehicles on the adjacent roads. The vehicles on the main road are divided into deceleration straight-ahead segments and lane-changing segments. The data of vehicles on the adjacent roads are divided into avoidance segments and non-avoidance segments. After data extraction and classification, the imbalance between different categories of data will affect the final accuracy of the model. The SMOTE method is used for data enhancement. The SMOTE method expands the minority class data by synthesizing minority class samples. In the feature space, this method generates new synthetic samples by interpolating minority class samples with their neighboring samples, thereby balancing the category distribution in the data set. First, identify the minority samples in the dataset and separate them from the dataset. For an unbalanced dataset, the number of minority samples is much smaller than that of majority samples. Use the K-Nearest Neighbors (KNN) algorithm to calculate the K nearest neighbor samples of each minority sample in the feature space. For each minority sample A, randomly select a neighbor sample from its K closest minority samples B1, B2, ..., BK, and generate a new synthetic sample between sample A and the neighbor sample. The synthetic sample is generated by performing a synthetic operation on A and the neighbor sample B. i Linearly interpolate the feature differences between: S=A+σ×(B i -A) (1) Where S represents the synthetic sample, σ is a random number between 0 and 1, which is used to control the position of interpolation; Repeat this process until a sufficient number of synthetic samples are generated, and add the generated synthetic samples to the original data set to form a new balanced data set so that the number of minority class samples reaches the expected balance target; the upper-level game theory model in the two-level decision model described in the second step is designed as follows: 1) Determine the game participants: the merging vehicle (MV) and the following vehicle (FV); 2) Determine the strategy set of game participants: S MV {Waiting for merging, merging}, S FV {speed up, maintain stability, and give way politely}; 3) Revenue function design: The design goal of MV is to complete the merger as quickly as possible while ensuring safety, and the design goal of FV is to ensure safety and minimize speed fluctuations; The FV profit function is designed as follows: When the MV chooses to merge and the FV chooses to yield politely, the predicted acceleration required by the FV is calculated as follows: Where w is the width of MV, d1 and d2 are the distances from FV to the left and right rear of MV, respectively, and s e is an intermediate variable, s* represents the ideal distance between FV and the preceding vehicle, s0 represents the minimum distance between FV and the preceding vehicle, T represents the ideal headway, and v F 、v M Respectively represent the speed of FV and MV at the current moment, a max represents the maximum acceleration of FV, d comfort represents the comfortable deceleration of FV, v0 represents the ideal speed of FV, Acc FV_md This is the predicted acceleration required for FV; At this time, the profit function of FV is: U FV_md =α1+λ1Acc FV_md (5) Where α1 and λ1 are the parameters to be calibrated; When MV chooses to merge and FV chooses to accelerate, the predicted acceleration required by FV is calculated as follows: The speed of the MV at the predicted merging time is calculated based on the speed and acceleration of the MV at the current moment and the remaining distance in the acceleration lane: In the formula, a M represents the acceleration of MV at the current decision moment, RD represents the remaining distance of MV in the acceleration lane; The remaining time of MV in the acceleration channel is calculated as follows: The predicted acceleration required for FV is calculated as follows: v' F =v F +a F t' M (8) Where v' F represents the speed of predicting FV, a F represents the acceleration of FV at the current decision moment, X and X' represent the distance between FV and MV at the current decision moment and the prediction moment, respectively, t b is the reaction time; At this time, the profit function of FV is: U FV_ma =α2+λ2Acc FV_ma (11) Where α2 and λ2 are the parameters to be calibrated; When the MV selections are merged and the FV selections remain unchanged, the predicted acceleration required by the FV is calculated as follows: As shown in formula (12), if the current MV speed is greater than the FV speed, the MV will complete the merger without interfering with the FV. If the MV speed is lower than the FV speed, the MV will force the FV to slow down. At this time, the profit function of FV is: And FV_mdn =α3+λ3Acc FV_mdn (13) Where α3 and λ3 are the parameters to be calibrated; When MV chooses to wait, the calculation of FV’s payoff function is similar; The profit function of MV is designed as follows: When the MV chooses to merge and the FV chooses to give way politely, the MV merges at a comfortable acceleration. The predicted acceleration required by the MV is: Acc MV_md =Acc comfort (14) Where Acc comfort It is the comfortable merge acceleration of MV; The profit function of MV is: The MV_md =β1+η1Acc MV_md (15) Where β1 and η1 are the parameters to be calibrated; When the MV chooses to merge and the FV chooses to accelerate or keep the same strategy, the MV needs to travel at the maximum acceleration to ensure that it reaches the merging point before the FV. The acceleration required by the MV is as follows: Acc MV_ma =Acc MV_mdn =Acc max (16) Where Acc max represents the maximum acceleration of the MV; At this time, the profit functions corresponding to MV are: U MV_ma =β2+η2Acc MV_ma (17) And MV_mdn =β3+η3Acc MV_mdn (18) Where β2, β3, η2, η3 are the parameters to be calibrated; When the MV chooses to wait and the FV chooses to give way politely, the MV merges with the FV at a comfortable acceleration after determining the FV avoidance strategy. The predicted acceleration required by the MV is: Acc MV_wd =Acc comfort (19) At this time, the profit function of MV is: U MV_wd =β4+η4Acc MV_wd (20) Where β4 and η4 are the parameters to be calibrated; When MV chooses to wait and FV chooses to accelerate or maintain the status quo, MV needs to wait for FV to overtake before merging. The waiting time of MV is calculated as follows: The predicted acceleration required by the MV at this time is calculated as follows: The profit function of MV is: U MV_wa =β5+η5Acc MV_wa (23) U MV_wdn =β6+η6Acc MV_wdn (24) Where β5, β6, η5, η6 are the parameters to be calibrated; The design steps of the lower-level game theory model in the two-level decision-making model are as follows: 1) Determine the game participants: the main lane vehicle (FV) and the adjacent lane rear vehicle (Lag); 2) Determine the strategy set of game participants: S FV {Change lane, go straight}, S Lag {slow down to avoid, do not avoid}; 3) Benefit function design: The design objectives of FV and Lag are to ensure safety and minimize speed fluctuations; The profit function design of FV and Lag is as follows: The deceleration of the vehicle to avoid a collision is: In the above formula, v B represents the speed of the vehicle ahead, v A Indicates the speed of the rear vehicle, s B represents the position of the vehicle ahead, s A Indicates the position of the rear vehicle, l A Indicates the length of the vehicle behind; The profit function of FV and Lag is calculated as follows: In the formula, i represents the decision of FV, j represents the decision of Lag, and a FV,ij represents the acceleration of FV when FV makes decision i and Lag makes decision j, a Lag,ij represents the acceleration of Lag when FV makes decision i and Lag makes decision j, μ1, μ2, μ3, μ4, ψ1, ψ2 and ψ3 are the parameters to be optimized; The upper-level game theory model and the lower-level game theory model constitute a two-level decision-making model. When MV enters the acceleration channel, the upper-level game theory matrix begins to be constructed. When FV chooses to give way politely after bargaining with MV, FV then bargains with Lag to decide whether to change lanes or slow down and go straight.

2. A method for establishing a high-speed entrance ramp forced merging decision model based on game theory as described in claim 1, characterized in that: The objective function for the calibration of the parameters of the upper-level game theory model in the third step is as follows: In the formula, F i and They represent the decision results of the actual data set FV and the decision results predicted by the model, respectively, and M i and They represent the decision results of the actual MV data set and the decision results predicted by the model, respectively. i represents the index of the event in the data set, and n represents that a total of n events are extracted from the data set. The objective function for parameter calibration of the underlying game theory model is as follows: Where, L i and They represent the decision results of the Lag actual data set and the decision results predicted by the model respectively, and m means that a total of m events are extracted from the data set.

3. A method for establishing a high-speed entrance ramp forced merging decision model based on game theory as described in claim 1, characterized in that: The confusion matrix described in step 4 contains many important performance indicators, including: True positive: The case where the decision predicted by the model is consistent with the decision of the actual data set, reflecting the correct prediction ability of the model; False Positive (FP): The decision result predicted by the model is inconsistent with the decision result of the actual data set, indicating that the model has misjudgment; Detection rate: the proportion of correctly predicted events among all actual events, which measures the sensitivity of the model; False alarm rate: the proportion of incorrectly predicted events to all events in the data set, indicating the false alarm frequency of the model; In addition, to gain a more comprehensive understanding of the performance of the proposed model, the model performance of each strategy was independently evaluated to further validate the accuracy of the model.

4. A device for establishing a high-speed entrance ramp forced merging decision model based on game theory, characterized in that: The method comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, the method is used to implement a method for establishing a high-speed entrance ramp forced merging decision model based on game theory as described in any one of claims 1 to 3.

5. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, a method for establishing a high-speed entrance ramp forced merging decision model based on game theory as described in any one of claims 1 to 3 is implemented.

Citation Information

Patent Citations

  • Ramp cooperative control system and method based on driving style

    CN114789729A