Method for identifying and analyzing vehicle active avoidance behavior based on trajectory data
By combining trajectory data and machine learning technology, identifying and analyzing vehicle's active avoidance behavior, the problem of failing to fully consider the factors of surrounding vehicles and drivers in existing research is solved, and the accurate identification and in-depth analysis of active avoidance behavior is achieved, providing a scientific basis for strategy formulation.
Patent Information
- Application Number
- CN202411072853.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-08-06
AI Technical Summary
The existing research on factors influencing factors of active avoidance behavior has not fully considered the influence of surrounding vehicles, driver's personal style and real-time driving status, and needs to be improved in terms of the classification and identification methods of active avoidance types.
By comprehensively applying statistical methods and machine learning technology, combining the vehicle's running trajectory data in real scenes, active avoidance behavior scenarios and types are divided, and active avoidance behavior recognition methods and result recognition methods are proposed based on trajectory data. Multiple Logit models and machine learning models (such as XGBoost, CatBoost, CNN) were used to study the impact of various factors on active avoidance behavior types and outcomes.
Accurate identification and in-depth analysis of vehicle active avoidance behavior, revealing the mechanism and influencing factors of avoidance behavior, and providing a scientific basis for the formulation of active avoidance strategies.
Smart Images

Figure CN118770201B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road traffic safety, and in particular to a method for identifying and analyzing vehicle active avoidance behaviors based on trajectory data. Background Art
[0002] In the research on influencing factors of driving behaviors, domestic and foreign scholars have conducted in-depth research based on various data such as traffic accidents, traffic violations, driving simulation experiments, vehicle trajectories, etc., combined with factors such as weather, road design, time, and space. The research mainly focuses on analyzing the influencing factors of the occurrence and severity of traffic accidents, violations, and conflicts under the influence of dangerous driving behaviors.
[0003] Active avoidance behavior is one of the typical driving behaviors. In the research on vehicle active avoidance behaviors, the avoidance types mainly include braking, lane changing, and coordination (lane changing + braking). However, in actual situations, there are cases where the speed of the target lane is higher than that of the avoidance vehicle. At this time, the vehicle may need to accelerate and change lanes to avoid risks. Previous research has not fully considered this behavior, and there is room for improvement in the classification and identification methods of active avoidance types.
[0004] Due to data limitations, the existing research on influencing factors of active avoidance behaviors mainly considers vehicle dynamics, vehicle braking force, and reaction time, etc., and fails to fully consider factors such as the influence of surrounding vehicles, the driver's personal style, and the real-time driving state. Trajectory data has been widely used in the current research in the field of traffic safety. It records the high-precision speed and position information of the current vehicle and surrounding vehicles. By calculation, the driving style and driving state of the vehicle can be divided, which can provide rich microscopic data for the identification and analysis of active avoidance behaviors.
[0005] In view of this, the present application provides a method for identifying and analyzing vehicle active avoidance behaviors based on trajectory data. Summary of the Invention
[0006] The present application provides a method for identifying and analyzing vehicle active avoidance behaviors based on trajectory data. This solution comprehensively uses statistical methods and combines the running trajectory data of vehicles in real scenarios to define vehicle active avoidance behaviors under normal environments. After preprocessing the trajectory data, a method for identifying active avoidance behaviors is proposed to identify the active avoidance behaviors in the trajectory data and conduct a preliminary analysis of their characteristics. Through machine learning and a multinomial Logit model, an analysis model of influencing factors of active avoidance behaviors and their results is constructed to deeply explore the occurrence mechanism and influencing factors of avoidance behaviors.
[0007] A method for identifying and analyzing vehicle active avoidance behaviors based on trajectory data, characterized in that
[0008] First, based on the vehicle's running trajectory data in the real scenario, divide the active avoidance behavior scenarios and define the types of active avoidance behaviors;
[0009] Secondly, based on the vehicle running trajectory data, propose an active avoidance behavior recognition method and an avoidance result recognition method based on vehicle trajectory data;
[0010] Then, establish a model based on machine learning and a fixed-parameter multinomial Logit model to study the influence of various factors on the types of vehicle active avoidance behaviors;
[0011] Finally, establish a fixed-parameter Logit model, a random-parameter Logit model, and a random-parameter Logit model considering the mean and variance to study the influence of various factors on the avoidance results.
[0012] In summary, the present application has the following beneficial effects:
[0013] Comprehensively apply statistical methods, combine the vehicle's running trajectory data in the real scenario, and define the vehicle's active avoidance behavior in the normal environment. After preprocessing the trajectory data, propose an identification method for active avoidance behaviors to identify the active avoidance behaviors in the trajectory data and conduct a preliminary analysis of their characteristics. Through machine learning and the multinomial Logit model, construct an analysis model of the influencing factors of active avoidance behaviors and their results to deeply explore the occurrence mechanism and influencing factors of avoidance behaviors. Description of the Drawings
[0014] Figure 1 is a schematic diagram of the research technical route of the present application;
[0015] Figure 2 is a schematic diagram of the active avoidance behavior scenario of the present application;
[0016] Figure 3 is a schematic diagram of the active avoidance behavior selection of the present application;
[0017] Figure 4 is a schematic diagram of the active avoidance behavior recognition process of the present application. Detailed Embodiments
[0018] The following will further describe in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0019] 1 Vehicle Active Avoidance Behavior Recognition and Analysis Method Based on Trajectory Data
[0020] This application is based on the operation trajectory data of vehicles in real scenarios. First, it divides the active avoidance behavior scenarios and defines the types of active avoidance behaviors; then, it proposes a method for identifying active avoidance behaviors based on vehicle trajectory data, and studies the types and frequency distribution characteristics of vehicle active avoidance behaviors; finally, it conducts research on the influencing factors of vehicle active avoidance behaviors to provide a basis for formulating active avoidance strategies.
[0021] 1.1 Types of Vehicle Active Avoidance Behaviors
[0022] 1.1.1 Definition of Types of Vehicle Active Avoidance Behaviors
[0023] According to the operation methods and behavior characteristics (as Figure 3 shown), the vehicle active avoidance behaviors are divided into the following types:
[0024] (1) Braking Behavior
[0025] Vehicles in the conflict area reduce their speeds through braking operations, increase the distance from the vehicle in front, and avoid the collision risk with the vehicle in front.
[0026] (2) Lane-changing Behavior
[0027] In the conflict area, when the driving conditions of the surrounding lanes permit, the vehicle changes lanes at an appropriate speed according to the vehicle condition to avoid the collision risk with the vehicle in front. According to the speed changes of the vehicle during the lane-changing process, the lane-changing avoidance behaviors of the vehicle are divided into three types: uniform-speed lane-changing, accelerating lane-changing, and braking lane-changing.
[0028] ① Uniform-speed lane-changing: The driver transfers the vehicle to the adjacent lane while keeping the speed fluctuating within a small range;
[0029] ② Accelerating lane-changing: When the speed is lower than that of the target lane, the vehicle increases its driving speed while ensuring safety to achieve a quick lane change and avoid collision;
[0030] ③ Decelerating lane-changing: This behavior combines the decelerating behavior and the lane-changing behavior. When the speed of the vehicle in front is higher than that of the target lane, the vehicle appropriately decelerates and conducts a lane-changing operation to ensure safe lane change while avoiding the risk of rear-end collision.
[0031] 1.2 Research on the Method and Characteristics of Identifying Vehicle Active Avoidance Behaviors
[0032] 1.2.1 Method for Identifying Vehicle Active Avoidance Behaviors
[0033] Vehicle active avoidance behaviors are divided into four types: braking avoidance, uniform-speed lane-changing avoidance, accelerating lane-changing avoidance, and decelerating lane-changing avoidance. According to the defined scope of vehicle active avoidance behaviors, an identification process for the types of active avoidance behaviors is formulated based on the trajectory data after data preprocessing (as Figure 4As shown), in order to determine the safety of the vehicle's current driving state, reference is made to the conclusions of previous studies (see 88-90 below):
[0034]
[88] Fancher PS,Bareket Z,Ervin R D.Human-centered design of an acc-with-braking and forward-crash-warning system[J].Vehicle System Dynamics,2001,36:203-223.
[0035]
[89] Lee K,Peng H.Evaluation ofautomotive forward collision warning and collision avoidance algorithms[J].Vehicle System Dynamics,2005,43:735-751.
[0036]
[90] Jeppsson H, M, Lubbe N.Real life safety benefits ofincreasing brake deceleration in car-to-pedestrian accidents:Simulation of Vacuum Emergency Braking[J].AccidentAnalysis&Prevention,2018,111:311-320.
[0037] When TTC is greater than 7s, the vehicle has no rear-end collision risk. When it is less than 7s, the driver will take braking action. When it is between 0.7-2s, the vehicle enters the danger zone. When it is less than 0.7s, the vehicle is extremely dangerous and difficult to intervene.
[0038] This application uses the vehicle safety index (TTC) = 7s as the threshold for determining the conflict area. The determination conditions for different types of active avoidance behaviors of different vehicles are as follows:
[0039] (1) Braking avoidance judgment conditions
[0040] In order to avoid the occurrence of deceleration in a short period of time caused by the driver's accidental braking, data errors, and other non-driver subjective intentions during the recognition process, when identifying the braking avoidance behavior, the time it takes for the vehicle's braking deceleration to gradually rise to the maximum value is taken, that is, Δt c =0.2s, as one of the judgment conditions.
[0041] (2) Lane-changing avoidance determination conditions
[0042] When the TTC is less than 7 s, if the vehicle lane number changes, it is considered that a lane-changing behavior occurs, and then the uniform speed, acceleration, and deceleration behaviors of the vehicle during the lane-changing operation are further determined.
[0043] Given that in real data, there are extremely few cases where the speed change amount Δv of the vehicle is constantly 0 during lane-changing, considering allowing the vehicle to have a certain floating range, in this study, the average change amount of the vehicle speed during lane-changing The frequency fitting distribution result of is combined with cluster analysis to determine the thresholds for distinguishing acceleration, uniform speed, and deceleration during the vehicle lane-changing process. Using the K-means clustering analysis method, the average change amount of the vehicle speed during lane-changing is respectively divided into different clusters between 3 and 20, and the optimal number of clustering clusters is screened as 15 by comprehensively evaluating the silhouette coefficient, Davies-Bouldin score, and Calinski-Harabasz index. The value of the 9th cluster is in the range of [-0.04, 0.00], corresponding to the value range [μ - 1 / 7σ, μ + 1 / 7σ), that is, [-0.041, 0.004). Therefore, according to the comparative analysis, during the vehicle lane-changing:[[]] When, it is vehicle acceleration; When, it is vehicle uniform speed; When, it is vehicle acceleration.
[0044] ① Uniform-speed lane-changing avoidance
[0045] When the vehicle has a lane-changing behavior within the conflict area, and at the same time the average value of the speed change amount
[0046] When, it is considered that the type of active avoidance behavior of the vehicle is uniform-speed lane-changing avoidance.
[0047] ② Acceleration lane-changing avoidance
[0048] When the vehicle has a lane-changing behavior within the conflict area, and at the same time the average value of the speed change amount When, it is considered that the type of active avoidance behavior of the vehicle is acceleration lane-changing avoidance.
[0049] ③ Braking lane-changing avoidance
[0050] When the vehicle has a lane-changing behavior within the conflict area, and at the same time the average value of the speed change amount When, it is considered that the type of active avoidance behavior of the vehicle is braking lane-changing avoidance.
[0051] Through the avoidance behavior recognition method based on trajectory data, a total of 8,215 active avoidance behaviors were identified (as shown in Table 1.1), specifically including 216 acceleration lane changes, 96 constant-speed lane changes, 242 deceleration lane changes, and 7,661 brakings.
[0052]
[0053] 1.2.2 Vehicle Active Avoidance Behavior Result Discrimination Method
[0054] To study the effect of active avoidance behavior, the avoidance result is defined as whether the time to collision (TTC) between the vehicle and the vehicle in front is greater than 7 s after a single active avoidance behavior of the vehicle. Analyzing the statistical characteristics of the avoidance results, it is found that among different types of active avoidance behaviors of the vehicle, after a single lane change avoidance behavior (acceleration lane change, constant-speed lane change, and deceleration lane change), the successful avoidance ratios of the three lane change avoidance behavior types are all above 71.76%. Among them, the proportion of constant-speed lane changes is as high as 86.46%, while the successful avoidance ratio of a single braking avoidance behavior is only 18.42%, fully indicating that lane change operations are more effective in avoiding rear-end collision risks (as shown in Table 1.2); the behavior of entering the comfort zone by a single braking avoidance only accounts for 18.43%, indicating that when the vehicle takes braking to avoid risks, multiple braking operations are usually required.
[0055] Table 1.2 Active Avoidance Behavior Result Statistics
[0056] Avoidance result Accelerated lane change Constant-speed lane change Decelerated lane change Braking Total Enter the comfort zone after a single avoidance 155 83 191 1412 1841 Do not enter the comfort zone after a single avoidance 61 13 51 6249 6374 Total 216 96 242 7661 8215
[0057] 1.3 Modeling and Result Analysis of Factors Affecting Vehicle Active Avoidance Behavior
[0058] 1.3.1 Research on Active Avoidance Behavior Types Based on Machine Learning and Multinomial Logit Model
[0059] Objectively, the types of vehicle active avoidance behaviors are affected by factors such as drivers, the vehicle itself, and the driving environment. Therefore, machine learning models (XGBoost, CatBoost, convolutional neural network) and a fixed-parameter multinomial Logit model are established to study the influence of various factors on the types of vehicle active avoidance behaviors. The characteristic variables representing the influencing factors in the model are shown in the following table:
[0060] Vehicle Candidate Feature Parameters and Definitions
[0061]
[0062] And indicators such as accuracy, recall rate, and F1 value are used to measure the performance of the model. In addition, the SHAP method based on the machine learning model is combined with the elasticity analysis results of the fixed-parameter multinomial Logit model to explain the influence of different characteristic variables in the model on the types of active avoidance behaviors.
[0063] (1) Data resampling
[0064] Among the identified active avoidance behaviors, the frequency of braking avoidance behavior is 7,661 (93.26%), and the frequency of uniform speed avoidance behavior is 96 (1.17%). The imbalance ratio reaches 30:1, presenting a serious data imbalance problem. The use of imbalanced data will cause the classifier to tend to classify the data into the majority class, resulting in problems such as overfitting and reduced model sensitivity, which will affect the overall performance of the model. Therefore, before building the model, the Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) methods are used to resample the data.
[0065] (2) Model construction
[0066] ① Machine learning model
[0067] The machine learning models adopt XGBoost (eXtreme Gradient Boosting), CatBoost (Categorical Boosting), and Convolutional Neural Network (CNN for short) models. XGBoost is an ensemble learning algorithm based on Gradient Boosting Tree. By continuously iteratively training multiple decision trees, the new decision tree fits the residuals between the current predicted value and the true label in each iteration. Improved methods such as regularization, tree pruning, and feature selection are introduced to control the complexity of the model, which is not easily affected by data anomalies and missing values and has good stability. CatBoost is also based on the Gradient Boosting Tree model framework. The difference is that CatBoost embeds a numerical feature algorithm and uses a sorting method to solve the problems of gradient estimation bias and prediction deviation. XGBoost constructs each tree based on a binary tree structure and uses a greedy algorithm to only consider one splitting method, which may lead to an asymmetric tree structure; the sorting-based gradient boosting method of CatBoost constructs decision trees based on a binary tree structure and selects the best splitting point from multiple splitting schemes considering the sorting information of feature values for classification, making the tree structure relatively symmetric and not prone to overfitting. When analyzing influencing factors based on the CNN model, the vehicle trajectory data is regarded as a one-dimensional image, and the CNN model is used to extract and learn the feature variables representing vehicle dynamics and traffic states to achieve the classification of active avoidance behavior types. CNN includes an input layer, multiple hidden layers, and an output layer. The hidden layers of CNN usually consist of a series of convolutional layers, pooling layers, and fully connected layers. The hyperparameters of the convolutional layer include the number of convolutional layers, the number of convolutional kernels, the kernel size, and the dilation rate. The hyperparameters such as the number of convolutional layers, the number of convolutional kernels, the kernel size, the dilation rate, the number of fully connected layers, the number of neurons, and the dropout rate in the CNN model are determined by Bayesian optimization, and the candidate set is shown in Table 1.3.
[0068] Table 1.3 Candidate Set of Hyperparameters for CNN
[0069] Hyperparameter Candidate set Number of convolutional layers [2,3,4] Number of convolutional kernels [32,64,128,256,512] Kernel size [3,4,5,6] Dilation rate [1,2,3,4,5] Number of fully connected layers [1,2,3] Number of neurons [256,512,1024] Dropout rate [0,0.1,0.2,0.3,0.4,0.5]
[0070] Due to the insufficient interpretability of machine learning, the SHAP algorithm is considered to explain the results of machine learning models. SHAP (Shapley Additive explanations) is a model interpretation tool proposed based on the cooperative game theory of Shaply values, used to explain the output of machine learning models, calculate the average of the marginal contributions of features to the model prediction results, and represent the contribution degree of each feature to the prediction results.
[0071] ② Multinomial Logit Model
[0072] 1) Collinearity test
[0073] Multicollinearity means there is a correlation between the characteristic variables of the model. The existence of multicollinearity will affect the accuracy of parameter estimation results. At the same time, excessive multicollinearity may lead to overfitting of the model, greatly reducing the model performance. When constructing the model, estimate the multicollinearity between the characteristic variables, and use the variance inflation factor (VIF) to test the correlation between the characteristic variables. When the maximum value of the variance inflation factor is greater than 10, it is considered that there is high multicollinearity. Input the characteristic variables into the model one by one, and at the same time observe the goodness of fit and the significance level of the variables, and retain the significant variables and the best fitting model. Since the categorical variables representing the same characteristic are correlated, select the categorical variables with the largest number among them for collinearity test.
[0074] The test steps are shown in Table 1.4.
[0075]
[0076]
[0077] 2) Model establishment
[0078] In the fixed-parameter multinomial Logit model, the vehicle active avoidance behavior type is the dependent variable, including four behaviors: decelerating and changing lanes to avoid, maintaining a constant speed and changing lanes to avoid, accelerating and changing lanes to avoid, and braking to avoid. The utility function of the influence of each characteristic variable on the active avoidance behavior type is shown in Equation (1.1).
[0079] U ij =β i X ij +ε ij (1.1)
[0080] In the formula, X ij is each characteristic variable; β i is the parameter vector of each characteristic variable; ε ij is the error term.
[0081] The expression of the multinomial Logit model is shown in Equation (1.2).
[0082]
[0083] In the formula, P ij is the probability of the active avoidance behavior type i in sample j; I is the set of all active avoidance behavior types.
[0084] Nlogit is a standard software package for evaluating discrete choice models, integrating all the functions of statistical programs. Therefore, the software Nlogit 5.0 is used to estimate the model parameters. Given that the maximum likelihood estimation of the random parameter logit model requires numerical integration of the logit formula and the calculation is complex, the maximum likelihood estimation based on simulation is adopted. Bhat et al. (Bhat C R. Simulation estimation of mixed discrete choice models using randomized and scrambled Halton sequences[J]. Transportation Research Part B: Methodological, 2003, 37: 837-855) found that using Halton sampling with the maximum likelihood method based on simulation is superior to pure random sampling. Therefore, the Halton sampling method is used to conduct 200 samplings to estimate the parameters instead of the random sampling method. To obtain the best model prediction performance, forward stepwise regression is adopted, that is, starting from the null model, each time a feature variable is added and the model is refitted, the evaluation index is calculated and the significance test of the independent variable is carried out, the insignificant feature variables are removed, and the significant variables are added to the model to determine whether the variable is a random variable and the distribution it follows one by one. At the same time, considering the possible distributions of the model parameters, including normal distribution, lognormal distribution and uniform distribution, when the standard error of the assumed distribution is not equal to zero at the 10% significance level, the corresponding parameter is considered random.
[0085] ③ Model evaluation method
[0086] 1) Accuracy is the ratio of the number of samples correctly classified by the classifier to the total number of samples, and is used to evaluate the overall performance of the model. The calculation formula is shown in Equation (1.3):
[0087]
[0088] In the formula, TP is the true positive, that is, the number of samples correctly predicted as positive by the classifier; TN is the true negative, that is, the number of samples correctly predicted as negative by the classifier; FP is the false positive, that is, the number of samples wrongly predicted as positive by the classifier; FN is the false negative, that is, the number of samples wrongly predicted as negative by the classifier.
[0089] 2) Recall is the ratio of the number of samples correctly predicted as positive by the classifier to the total number of actual positive samples.
[0090]
[0091] 3) The F1 value is the harmonic mean of precision and recall.
[0092]
[0093] In the formula, Precision is the precision.
[0094] The formula for calculating precision is as follows:
[0095]
[0096] (3) Model evaluation
[0097] ① XGBoost
[0098] An XGBoost model is constructed based on the original data and the resampled data, where 80% is used for model training and 20% is used as test samples. The test results are shown in Table 1.5. The results show that the XGBoost model constructed with unbalanced data cannot successfully estimate the behavior categories with small data volumes such as decelerating lane changes, constant-speed lane changes, and accelerating lane changes, and there is a significant gap compared with braking avoidance. Visualize the F1 values of different avoidance types. The F1 value of constant-speed lane change in the original data is only 0.15. The XGBoost models established based on the two resampling methods of SMOTE and ADASYN both ensure the estimation accuracy of braking avoidance while making up for the estimation deviation of the active avoidance behavior types with small data volumes. Among them, the XGBoost model constructed by resampling data based on the SMOTE method correctly estimated 4537 samples out of 4900 test samples, with a correct estimation rate of 92.59%, while the XGBoost model constructed by resampling data based on the ADASYN method correctly estimated 93.23%, indicating that the XGBoost model constructed by resampling data based on the ADASYN method has a lower misclassification rate.
[0099] Table 1.5 Comparison of XGBoost model results under different resampling methods
[0100]
[0101] ② CatBoost
[0102] Construct a CatBoost model with the same test ratio and hyperparameter settings. The results are shown in Table 1.6. By studying the confusion matrix, it is found that the estimation of the CatBoost model constructed based on the original data has a serious deviation. The prediction accuracies of decelerating lane change and constant-speed lane change are 0, and the model is more inclined to estimate the samples as braking and avoiding behaviors. The prediction accuracies of the CatBoost models constructed after resampling by the SMOTE and ADASYN methods are similar. Combining the F1 value results, it can be seen that the prediction accuracy of the CatBoost model is lower than that of the XGBoost model, and XGBoost shows better performance through different resampling techniques.
[0103]
[0104] ③ CNN model
[0105] The optimized CNN model has 4 convolutional layers and 2 fully connected layers. In each convolutional layer, the number of convolutional kernels, kernel size, and dilation rate are 256, 3, and 3 respectively. In each fully connected layer, the number of neurons and dropout rate are 512 and 0.2 respectively. The estimation results are shown in Table 1.7. By studying the confusion matrix, it is found that the CNN model constructed based on the original data also has a serious estimation deviation, and the model is more inclined to estimate the samples as braking and avoiding behaviors. The CNN models constructed based on the two resampled data show relatively high prediction accuracies. Combining the F1 value results, it can be seen that the prediction accuracy of the CNN model is higher than that of the CatBoost model and the XGBoost model, and the F1 values are all above 0.90. Among them, for the model constructed by the ADASYN resampling method, the FI values are all higher than 0.95.
[0106] Table 1.7 Comparison of CNN model results under different resampling methods
[0107]
[0108]
[0109] ④ Fixed-parameter multinomial Logit model
[0110] Table 1.8 Comparison of multinomial Logit model results under different resampling methods
[0111]
[0112] The imbalanced data makes the four models constructed tend to estimate the majority class of braking and avoiding, with a low prediction level (as shown in Table 1.8). Compared with the CNN model, the F1 values of the minority samples are relatively low, reflecting the deficiency of traditional statistical models in estimating minority samples.
[0113] The F1 value of the model constructed by resampling data with SMOTE for predicting lane-changing avoidance behavior is higher than that of the ADASYN model. Although the prediction performance is improved after resampling by both SMOTE and ADASYN methods, it is found by observing the confusion matrix that the proportion of incorrect predictions is still higher than that of the three machine learning models.
[0114] Based on the above conclusions, a further study is carried out using the fixed-parameter multinomial Logit model constructed with the data resampled by the SMOTE method.
[0115] (4) Model result analysis
[0116] Comparing the results of the machine learning models, the multinomial Logit model with better performance is selected to analyze the influencing factors of active avoidance behavior. The final estimation results of the multinomial Logit model are shown in Table 1.9.
[0117] Comparing the significance of the influencing factors of different active avoidance types, the results show that: the speed of the target vehicle only has a significant impact on the accelerating lane-changing behavior; the speed and acceleration of the vehicle behind have no significant impact on the accelerating lane-changing behavior. Relative to the vehicle behind, the vehicle pays more attention to the relevant factors of the vehicle in front when avoiding risks through lane-changing; except for the impact of the speed of the vehicle in the right rear on braking avoidance, the relevant factors of other adjacent lanes all significantly affect the choice of the vehicle's active avoidance behavior, indicating that the vehicle fully considers the traffic conditions of adjacent lanes when performing active avoidance behavior; an aggressive driving style has a significant impact on braking and decelerating lane-changing behaviors.
[0118] The feature variables are divided into four categories: vehicle itself related factors, vehicle in front related factors, vehicle behind related factors, and adjacent lane related factors to further analyze the model results. To study the impact of changes in different feature variables on the types of vehicle active avoidance behaviors, the average probability elasticity of each category of factors is calculated to represent the impact of the occurrence of each feature variable on the possibility of vehicle active avoidance behavior selection.
[0119] Table 1.9 Influence of feature variables on active avoidance behavior types (multinomial Logit model)
[0120]
[0121]
[0122] Note: ***, **, * ==> Significance at the 1%, 5%, 10% levels.
[0123] 1.3.2 Modeling of active avoidance behavior results considering heterogeneity and analysis of influencing factors
[0124] There are differences or diversities, that is, heterogeneities, among and within individuals in a certain group or system, including heterogeneities among individuals, heterogeneities within individuals, and heterogeneities of error terms, etc. The Logit model assumes a log-linear relationship between feature variables and the dependent variable, but there is often a non-linear relationship between feature variables and the dependent variable, and it is difficult to collect and consider all feature factors in the research. Therefore, there is unobserved heterogeneity in the traditional Logit model. If the unobserved heterogeneity is ignored, the estimated parameters are usually biased and inefficient, which will lead to incorrect inferences and predictions. To study the effect of active avoidance behavior, using the avoidance result data, that is, whether the vehicle can enter the comfortable area after a single active avoidance behavior, a fixed parameter model, a random parameter model, and a random parameter model considering the mean and variance are constructed to analyze the influence of each feature variable on the result of active avoidance behavior. Among them, the random parameter Logit model considers the heterogeneity of error terms and individuals, and the random parameter Logit model of mean and variance heterogeneity considers the heterogeneity within individuals while considering the heterogeneity of error terms and individuals.
[0125] (1) Model establishment
[0126] ① Fixed parameter Logit model
[0127] The stepwise regression method is used to select key variables. Based on the vehicle active avoidance result data, the final binary Logit model is gradually obtained. The linear utility function (Equation 1.7) of different avoidance results i is introduced, where the error term is assumed to follow a certain generalized extreme value distribution. If it follows the logistic distribution, the probability of successful vehicle avoidance is deduced, as shown in Equation (1.8).
[0128] U ij =α i +β i X ij +ε ij (1.7)
[0129] In the formula, α i is the constant term; X ij are the feature variables; β i ' is the parameter vector of each feature variable; ε ij is the error term.
[0130]
[0131] In the formula, P y=i is the probability of the active avoidance behavior result i.
[0132] ② Random parameter Logit model
[0133] The Random Parameters Logit Model is an extension of the Logit model, which is used to account for unobserved heterogeneity. In the Random Parameters Logit Model, each parameter βi is assumed to be a random parameter β randomly drawn from a certain probability distribution in , to account for heterogeneity among individuals, as shown in Equation (1.9), where the mean and variance of β in are fixed values.
[0134]
[0135] In the formula, β is the average parameter estimate of all observations; is the random distribution term.
[0136] ③ Random Parameters Logit Model with Mean and Variance Heterogeneity
[0137] Compared with the non-heterogeneous random parameters model, the random parameters model with heterogeneous means and variances improves the analysis accuracy and explanatory power of the model. The random parameters Logit model with heterogeneous means and variances assumes that the parameters of the distribution equation followed by the model parameters are random and follow a certain probability distribution. The model takes into account the heterogeneity within individuals while considering the heterogeneity among individuals.
[0138] In the model, the means and variances of each random parameter β in are not fixed. As shown in Equation (1.10), if Z in or w i is not statistically significant, then there is no heterogeneity in the mean or variance of the model random parameters.
[0139]
[0140] In the formula, δ in is the corresponding vector of the estimable parameters; Z in is used to capture the heterogeneity of the random parameter mean; σ in is the quasi-deviation; Ψ i is the estimable parameter vector; w i is the heterogeneity parameter for capturing the quasi-deviation σ in .
[0141] (2) Model Evaluation
[0142] The goodness of fit of the model is the degree to which the model fits the data. The complexity of the model refers to the number of parameters included in the model. A complex model can fit the data better, but it is prone to overfitting problems.
[0143] 1) The Akaike information criterion (AIC) is a statistical metric for model selection that takes into account both the goodness of fit and complexity of the model. For models with different complexities, the smaller the AIC value, the better the goodness of fit of the model. The calculation formula is shown in Equation (1.11):
[0144] AIC = -2ln(L) + 2K (1.11)
[0145] Where L is the maximum likelihood estimate of the model; K is the number of parameters.
[0146] In the Akaike information criterion, the maximum likelihood estimate measures the goodness of fit of the model, and the complexity of the model serves as a penalty term for the metric. Different models have different degrees of fit and complexity. When using AIC to evaluate model performance, consider choosing the model with the smallest AIC value.
[0147] 2) Bayesian information criterion
[0148] The Bayesian information criterion (BIC) is another model selection metric similar to AIC that aims to balance the goodness of fit and complexity of the model and is used to compare the performance of multiple models. The smaller the BIC value, the better the goodness of fit of the model. The calculation formula for BIC is as follows:
[0149] BIC = -2ln(L) + ln(n)×K (1.12)
[0150] Where L is the maximum likelihood estimate; K is the number of parameters; n is the sample size of the data set.
[0151] Different from the Akaike information criterion, the penalty term of the Bayesian information criterion is more stringent, taking into account both the number of parameters and the data sample size. In this study, the accuracy, sensitivity, and specificity of model performance evaluation were considered. It should be noted that previous studies have also used the false positive rate to measure predictability.
[0152] 3) The likelihood ratio test (LRT) is based on the maximum value of the likelihood function to compare the goodness of fit of two models with different levels of complexity. By calculating the maximum values of the likelihood functions of two different models, the log-likelihood function values of the simple model and the complex model are denoted as LL(A) and LL(B) respectively. Then the likelihood ratio test statistic is shown in Equation (1.13):
[0153] X2 = -2[LL(A) - LL(B)] (1.13)
[0154] The result value X2 follows a chi-squared distribution (with the degrees of freedom being the difference between the significant variables of the two models). Using the degrees of freedom of the likelihood ratio and the significance level p-value (taking 0.01), the corresponding critical value of the chi-squared distribution is found. The likelihood ratio is compared with this critical value. If the critical value is greater than the likelihood ratio, the p-value is less than 0.01, indicating that the difference between the two models has greater statistical significance, and the simple model is rejected. If the critical value is less than or equal to the likelihood ratio, the p-value is greater than 0.01, and the difference between the two models is not significant.
[0155] (3) Model comparison
[0156] The fitting metric results of the three models for analyzing the influencing factors of the active avoidance behavior results are shown in Table 1.10. The performance of the random parameter model in the heterogeneity models of the mean and variance is compared with that of the fixed parameter model and the random parameter model. The results show that: from the fixed parameter model, the random parameter model to the random parameter model with mean and variance heterogeneity, the AIC and BIC values decrease in turn, and the random parameter model with mean and variance heterogeneity has a higher log-likelihood function value; the likelihood ratio test values of the random parameter model with mean and variance heterogeneity and the other two models are 69.96 and 122.41 respectively, both significant at the 0.001 level. Therefore, the model considering mean and variance heterogeneity has the highest fitting performance.
[0157] Table 1.10 Fitting metrics of the active avoidance behavior result estimation model
[0158] Model Fixed-parameter model Random-parameter model Heterogeneity model Number of parameters 11 14 17 Sample size 8215 8215 8215 Log-likelihood function value -3767.92 -3692.95 -3576.01 AIC 7557.80 7423.90 7185.1 BIC 7635.00 7522.10 7304.3 X2 69.96 122.41 Degree of freedom 3 3 Significance level <0.001 <0.001
[0159] (4) Model result analysis
[0160] The estimation results of the Logit model considering mean and variance heterogeneity, which shows the best performance, are shown in Table 1.11. In addition, the marginal effects of each influencing factor are estimated in the study, and the following conclusions are drawn based on the results.
[0161]
[0162]
[0163] Note: ***, **, * ==> Significance at the 1%, 5%, 10% levels.
[0164] This application divides the scenarios where active avoidance behaviors occur and defines the types of active avoidance behaviors of vehicles. It proposes a method for identifying active avoidance behaviors based on trajectory data and a method for discriminating active avoidance results. An analysis model of the influencing factors of active avoidance behavior types is constructed. In view of the data imbalance problem, two methods, SMOTE and ADASYN, are used to resample the data, and three machine learning models, XGBoost, CatBoost, and CNN, and a multinomial Logit model are constructed. The results of the CNN model with the highest prediction accuracy and the multinomial Logit model are comprehensively considered to deeply study the influencing factors of active avoidance behavior types. Finally, to more comprehensively analyze the results of active avoidance behaviors and their influencing factors, considering the heterogeneity of the mean and variance, a fixed parameter Logit model, a random parameter Logit model, and a random parameter Logit model considering the mean and variance are constructed for the results of vehicle active avoidance behaviors.
[0165] Through the method for identifying active avoidance behaviors based on trajectory data, a total of 8,215 active avoidance behaviors are identified. Among them, there are 216 cases of accelerating lane changes, 96 cases of constant-speed lane changes, 242 cases of decelerating lane changes, and 7,661 cases of braking. The research on the distribution characteristics of TTC and vehicle spacing of active avoidance behaviors shows that: (1) When TTC is relatively large, vehicles are more inclined to take lane-changing avoidance; (2) The frequency distribution of TTC is relatively discrete in lane-changing avoidance behaviors, while it is relatively concentrated in braking avoidance behaviors; (3) The frequency distribution of vehicle spacing is relatively discrete in active lane-changing avoidance behaviors. The research results of the influencing factors of active avoidance behavior types show that: (1) For every 1% increase in vehicle speed, the probability of choosing an accelerating lane change behavior increases by 0.76%, that is, when the vehicle is driving at a relatively high speed in the conflict area, it is more inclined to choose an accelerating lane change; (2) Aggressive driving styles often lead to risky operations such as braking, rapid steering, and dangerous lane changes, thus increasing the probability of the vehicle choosing accelerating lane change and braking behaviors; (3) For every 1% increase in TTC, the probability of decelerating lane change avoidance increases by 0.06%; for every 1% increase in vehicle spacing, the probabilities of constant-speed lane change and accelerating lane change increase by 0.04% and 0.60% respectively; as TTC and vehicle spacing increase, the probabilities of the vehicle choosing constant-speed lane change and accelerating lane change also increase. The research results of the influencing factors of active avoidance behavior results show that: (1) Vehicle spacing is a random parameter that generates mean heterogeneity, while vehicle speed, leading vehicle speed, and lane type are random parameters that generate variance heterogeneity; (2) There is a positive correlation between TTC, vehicle spacing, and successful avoidance; (3) The marginal effect value of the influence of the middle left lane on vehicle successful avoidance is 0.04, indicating that on the left middle lane, the probability of successful active avoidance behavior of the vehicle is relatively high.
[0166] Although the present invention has been illustrated and described with respect to preferred embodiments, those skilled in the art should understand that various changes and modifications can be made to the present invention as long as they do not exceed the scope defined by the claims of the present invention.
Claims
1. A vehicle active avoidance behavior recognition and analysis method based on trajectory data, characterized in that: First, based on the running trajectory data of the vehicle in the real scene, the active avoidance behavior scenarios are divided and the active avoidance behavior types are defined; Secondly, based on the vehicle running trajectory data, an active avoidance behavior recognition method based on vehicle trajectory data is proposed, and the type and frequency distribution characteristics of vehicle active avoidance behavior are studied; Then, a multinomial Logit model based on machine learning and fixed parameters was established to study the impact of various factors on the type of active avoidance behavior of the vehicle; Finally, a fixed parameter Logit model, a random parameter Logit model and a random parameter Logit model considering mean and variance were established to study the impact of various factors on the avoidance results.
2. The method for identifying and analyzing active vehicle avoidance behavior based on trajectory data according to claim 1, characterized in that: The vehicle running trajectory data processing and analysis includes vehicle running parameter calculation, vehicle braking and lane changing behavior speed characteristics research in the conflict area; Vehicle operation parameter calculation includes: outlier Savitzky-Golay smoothing, vehicle operation status and safety parameter calculation, and vehicle driving style classification; The study on the speed characteristics of vehicle braking and lane changing behavior in the conflict zone includes: the study on the acceleration and deceleration characteristics when the vehicle changes lanes, and the study on the deceleration characteristics when the vehicle brakes.
3. The vehicle active avoidance behavior recognition and analysis method based on trajectory data according to claim 1 is characterized in that: The vehicle active avoidance behavior scenario is divided into a middle lane, a left side lane and a right side lane.
4. The method for identifying and analyzing active vehicle avoidance behavior based on trajectory data according to claim 1, characterized in that: The vehicle active avoidance behavior types include braking behavior and lane changing behavior; The lane changing behaviors include uniform speed lane changing, accelerated lane changing, and decelerated lane changing.
5. The method for identifying and analyzing active vehicle avoidance behavior based on trajectory data according to claim 4 is characterized in that: According to the definition scope of the vehicle's active avoidance behavior, a process for identifying the type of active avoidance behavior is developed to determine the vehicle's active avoidance behavior.
6. The method for identifying and analyzing active vehicle avoidance behavior based on trajectory data according to claim 1, characterized in that: The machine learning adopts XGBoost model, CatBoost model and CNN model; In the fixed parameter multinomial Logit model, the vehicle active avoidance behavior type is the dependent variable, including deceleration lane change avoidance, constant speed lane change avoidance, acceleration lane change avoidance and braking avoidance. The utility function of each characteristic variable on the active avoidance behavior type is as follows: U ij =b i X ij +e ij In the formula, X ij are the characteristic variables; β i is the parameter vector of each characteristic variable; ε ij is the error term; The fixed parameter multinomial Logit model expression is as follows: Where P ij is the probability of active avoidance behavior type i in sample j; I is the set of all active avoidance behavior types.
7. The vehicle active avoidance behavior recognition and analysis method based on trajectory data according to claim 6 is characterized by: The model evaluation indicators of the machine learning include: accuracy, recall, and the harmonic mean of accuracy and recall; Among them, accuracy is the ratio of the number of samples correctly classified by the classifier to the total number of samples, which is used to evaluate the overall performance of the model. Its expression is: Where TP is the true positive example, that is, the number of samples correctly predicted as positive examples by the classifier; TN is the true negative example, that is, the number of samples correctly predicted as negative examples by the classifier; FP is the false positive example, that is, the number of samples incorrectly predicted as positive examples by the classifier; FN is the false negative example, that is, the number of samples incorrectly predicted as negative examples by the classifier; Recall rate refers to the ratio of the number of samples correctly predicted as positive examples by the classifier to the total number of actual positive examples, and its expression is: The F1 value is the harmonic mean of precision and recall, and its expression is: Where Precision is the accuracy, and its calculation formula is:
8. The method for identifying and analyzing active vehicle avoidance behavior based on trajectory data according to claim 1, characterized in that: A fixed parameter Logit model, a random parameter Logit model and a random parameter Logit model considering mean and variance were constructed to analyze the impact of each characteristic variable on the results of active avoidance behavior.
9. The vehicle active avoidance behavior recognition and analysis method based on trajectory data according to claim 8 is characterized by: The fixed parameter Logit model construction includes selecting key variables by stepwise regression method, and gradually obtaining the final binary Logit model based on the vehicle active avoidance result data, and introducing the linear utility function expression of different avoidance results i as follows: U ij =a i +b i ′X ij +e ij In the formula, α i is a constant term; X ij are the characteristic variables; β i ' is the parameter vector of each characteristic variable; ε ij is the error term; the error term is assumed to obey a generalized extreme value distribution. If it obeys the logistic distribution, the probability of successful vehicle avoidance is derived as follows: Where P y=i is the probability of active avoidance behavior outcome i, β i is the parameter estimate of the factors affecting the active avoidance result, X i To actively avoid factors that affect the results; The random parameter Logit model expression is: In the formula, β in is the parameter β in the random parameter Logit model i Assume that is a random parameter randomly drawn from a probability distribution, and β is the average parameter estimate of all observations; is a randomly distributed term; The random parameter Logit model expression for the mean and variance heterogeneity is: In the formula, δ in is the corresponding vector of estimable parameters; Z in is used to capture the heterogeneity of the mean value of the random parameter; σ in is the standard deviation; i is an estimable parameter vector; w i To capture the quasi-deviation σ in Heterogeneity parameter.
10. The vehicle active avoidance behavior recognition and analysis method based on trajectory data according to claim 9, characterized in that: The model fitting metric evaluation indicators of the machine learning include Akaike information criterion, Bayesian information criterion, and likelihood ratio test; Among them, the calculation formula of Akaike information criterion is: AIC=-2ln(L)+2K Where L is the maximum likelihood estimate of the model; K is the number of parameters; The calculation formula of Bayesian information criterion is: BIC=-2ln(L)+ln(n)×K Where L is the maximum likelihood estimate; K is the number of parameters; n is the sample size of the data set; The likelihood ratio test statistic expression is: X 2 =-2[LL(A)-LL(B)] Wherein, the maximum values of the likelihood function of the simple model and the complex model are denoted as LL(A) and LL(B), respectively.
Citation Information
Patent Citations
Vehicle avoidance method, device, computer equipment and storage medium
CN113619574A
Pilotless automobile active collision avoidance algorithm based on pedestrian trajectory prediction
CN115230690A