Driving delay degree multi-dimensional feature analysis system for urban traffic accident site

By designing a multi-dimensional feature analysis system, using the random parameter Logit model and the potential category and machine learning double-layer model, combined with the SHAP model, the driving delay problem caused by urban traffic accidents is solved, and effective analysis and prediction of various influencing factors is achieved, helping the traffic management department to optimize management decisions.

CN119964359AActive Publication Date: 2025-05-09ZHEJIANG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411903536.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-09
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The driving delay caused by urban traffic accidents is complex and involves a variety of influencing factors. It is difficult to effectively analyze and predict the existing technology.

Method used

A multi-dimensional feature analysis system for driving delay degree at urban traffic accident sites was designed, and comprehensive analysis was carried out using pre-processing module, individual heterogeneity analysis based on random parameter Logit model, potential category and machine learning double-layer model, and SHAP model.

Benefits of technology

Through multi-dimensional feature analysis, the degree of driving delay caused by urban traffic accidents can be systematically modeled and predicted, and traffic management departments can take preventive measures and emergency responses, optimize road design and management decisions, and reduce traffic congestion and pollution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964359A_ABST
    Figure CN119964359A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of traffic safety analysis, provides a driving delay degree multi-dimensional feature analysis system of an urban traffic accident site, and aims at extracting road dimensional features when an accident occurs based on driving delay data of the urban traffic accident site, and providing a targeted data filling scheme in combination with the characteristics of missing features. And according to statistical distribution, redundant feature screening is completed, and structured traffic modeling data of four dimension features of weather, road, POI and time is constructed. The traffic delay degree caused by urban traffic accidents is systematically modeled and analyzed from the three aspects of individual heterogeneity, group transferability and group heterogeneity, and the influence effect of significant variables in multi-dimensional influence factors is explored, so that a traffic control department is helped to take preventive measures and emergency response countermeasures; and a theoretical basis is provided for design optimization and management decision of the road, so that the influence caused by traffic accidents is relieved, the operation risk level of an accident road section is improved, secondary accidents are avoided, and traffic pollution is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of traffic safety analysis, and in particular to a multi-dimensional feature analysis system for the degree of traffic delay at a location of an urban traffic accident. Background Art

[0002] Traffic accidents not only cause property damage and casualties, but also lead to problems such as lane encroachment at the accident site, large speed differences between upstream and downstream vehicles, and complicated driving behaviors when vehicles enter and exit. [3] , causing sporadic traffic congestion, causing traffic delays, and reducing road capacity and operation management levels. Traffic delays at the accident site will further increase the possibility of consecutive accidents on the accident section, causing secondary traffic accidents. Therefore, it is necessary to analyze the traffic delays caused by urban traffic accidents at the accident site. By mining the potential information and distribution patterns of traffic data, the relationship between traffic accidents, duration, traffic delays and different influencing factors is explored. In actual scenarios, there are many factors related to traffic delays at traffic accident sites, and large-scale traffic data containing multiple influencing factors is required for analysis.

[0003] In the analysis of traffic big data models, due to the development of computer technology, the processing of traffic big data has become more efficient and faster; the models are also more diversified, and can utilize statistical models such as random parameter Logit, machine learning models such as random forests, and SHAP analysis models, increasing the feasibility of individual and group-level research on a large number of characteristics, and the heterogeneity and transferability analysis of data. Summary of the invention

[0004] In order to achieve the above object, the present invention provides the following technical solutions:

[0005] The present invention provides a multi-dimensional characteristic analysis system for the degree of traffic delay at a location of an urban traffic accident, the analysis system comprising:

[0006] A preprocessing module is used to preprocess the traffic delay degree dataset of urban accident sites;

[0007] The random parameter logit model (RPLHM) based on mean heterogeneity is used to analyze the individual heterogeneity of the degree of traffic delay at urban traffic accident sites.

[0008] Analysis; that is, the differences or instability of influencing factors in different samples. The differences in some characteristics of the samples lead to different effects of the same influencing factor on different samples. Find out the individual heterogeneous variables and some sample characteristics that cause individual heterogeneity; Influencing factor analysis, that is, combining the average marginal effect value, analyze the variables that have a significant impact on traffic delays and their impact differences from four dimensions: road, weather, points of interest and time;

[0009] Based on the random parameter logit model (RPLHM), the overall and local transferability analysis of artificially divided groups is carried out by combining with the transferability theory, that is, whether the influencing mechanisms of the two groups are similar, and whether the model of a single group can maintain good performance on the data of another group; the overall data set is divided into day / night and weekdays / holidays, and the consistency or difference of the influencing factors of traffic delays between groups is detected.

[0010] The latent class and machine learning two-layer model is used to identify, predict and analyze the degree of traffic delay in urban traffic accident sites for heterogeneous groups. Group identification is used to clarify the best scenario for heterogeneous group division; delay degree prediction is used to select the best machine learning model to predict the degree of delay for heterogeneous groups; group heterogeneity analysis is used to explain and analyze the global importance of different features in heterogeneous groups in combination with the SHAP model.

[0011] Furthermore, the specific analysis steps of the preprocessing model are:

[0012] (11) Based on the delay duration, the dependent variable, the degree of traffic delay, is divided into minor delay, severe delay, and extreme delay. Factors that have no explanatory significance and belong to the accident lag variables are preliminarily excluded, including variables such as the accident end time and the starting and ending longitude and latitude.

[0013] (12) Based on the accident description field and the urban road location, the road status, road type and road section characteristic information at the time of the accident are extracted to obtain derived variables including four categories: road status, road type, peak time and season;

[0014] (13) According to the characteristics of statistical distribution, the wind chill temperature and rainfall variable features with too high data loss rate and similar meanings to other variables were deleted; the variance values ​​of all variables were calculated, and the variables with a variance of 0 were deleted, that is, the variables with a single distribution and no discrimination, including speed bumps, yield signs, roundabouts, turning loops, traffic lights and deceleration; weather conditions, wind direction, air pressure, visibility, relative humidity, humidity and wind speed were filled;

[0015] (14) Discretize and bin the temperature, visibility, and wind speed variables, and convert all subcategories of weather conditions, wind direction, peak hours, seasons, temperature, visibility, and wind speed into independent binary features, using 0 and 1 to represent binary state variables;

[0016] (15) Based on the tolerance and variance inflation factor indicators of the regression model, collinearity diagnosis is performed on all variables in the road dimension, point of interest (POI) dimension, weather dimension, and time dimension to avoid the impact of excessive collinearity levels on the significance level and convergence of the model. The evaluation indicators are tolerance and variance inflation factor (VIF).

[0017] The tolerance is calculated as follows:

[0018]

[0019] In the formula, To make the X i The variables of delay degree are taken as dependent variables, and the remaining k-1 variables are taken as independent variables to establish the determination coefficient of linear regression model; the tolerance value range is between 0 and 1, and the closer it is to 1, the lower the degree of collinearity between independent variables; it is generally believed that when the tolerance is greater than 0.1, there is no multicollinearity between variables, and this paper also selects 0.1 as the indicator.

[0020] The variance inflation factor (VIF) can reflect whether there is multicollinearity between the observed values ​​of the independent variables and is calculated as follows:

[0021]

[0022] In the formula, VIF is the inverse of tolerance, and variable X i of and There is a positive correlation, The larger the value is, the stronger the correlation between the variable and other variables is. The criteria for judging collinearity are:

[0023]

[0024]

[0025] When standard ① is met, it means that there is multicollinearity among the variables; when standard ① and standard ② are met at the same time, it means that the data has serious multicollinearity.

[0026] Furthermore, in step (13), the method for filling in the weather conditions, wind direction, air pressure, visibility, relative humidity, humidity and wind speed is:

[0027] Cold platform data and mode were used to fill missing values ​​for weather conditions and wind direction variables;

[0028] The Missforest multiple interpolation filling method based on the random forest model is used to perform regression prediction filling for the air pressure, visibility, temperature, relative humidity and wind speed variables. The specific steps of filling are:

[0029] First, combine P variables with n data to be filled into X = (X1, X2, X3, ..., X p ) of n*P dimensional data, use the mode or mean for preliminary filling;

[0030] Secondly, sort the variables with missing rates according to their missing rates and select the variable X with the smallest missing rate. S As the primary fill variable;

[0031] Furthermore, with variable X S As the predictor variable, the remaining variables are used as feature variables and are divided into X according to whether they contain missing values. S Observed value X S Missing values X S Observations other than With X S The remaining observations other than the missing values

[0032] Finally, and Bring it into the random forest model for training, using Data to predict missing data

[0033] During multiple iterations of the random forest model, when the latest filling result changes from the previous filling result by X S When it is infinitely close to 0, the calculation stops and the missing value filling is completed.

[0034]

[0035] In the formula, and Respectively represent X S The latest and last filled data sets of the variable.

[0036] Furthermore, in step (14), the subcategories of weather conditions, wind direction, peak hours, seasons, temperature, visibility, and wind speed are all converted into independent binary features, and the method is as follows:

[0037] A multi-class variable is divided into Y1, Y2, ...Y i ...,Y d When there are d separate sub-categorical variables, a sub-category Y d It is set as a reference variable and will not be included in the modeling variable range. Then, the binary variable state is represented by 0 and 1. When the variable value is 1, it means that the variable state exists. In addition, in the multi-classification variable, if a subclass Y d When Y is a reference variable, when the values ​​of other subclasses are 0, d The state indicated is established.

[0038] Furthermore, in step (15), the evaluation indicators of collinearity diagnosis are tolerance and variance inflation factor (VIF);

[0039] The tolerance is calculated as follows:

[0040]

[0041] In the formula, To make the X i The variables of the delay degree are taken as dependent variables, and the remaining k-1 variables are taken as independent variables to establish the determination coefficient of the linear regression model; the tolerance value range is between 0 and 1, and the closer it is to 1, the lower the degree of collinearity between the independent variables; it is generally believed that when the tolerance is greater than 0.1, there is no multicollinearity between the variables, and this paper also selects 0.1 as the indicator;

[0042] The variance inflation factor (VIF) can reflect whether there is multicollinearity between the observed values ​​of the independent variables and is calculated as follows:

[0043]

[0044] In the formula, VIF is the inverse of tolerance, and variable X i of and There is a positive correlation, The larger it is, the stronger the correlation between the variable and other variables;

[0045] The criteria for judging collinearity are:

[0046]

[0047]

[0048] When standard ① is met, it means that there is multicollinearity among the variables; when standard ① and standard ② are met at the same time, it means that the data has serious multicollinearity.

[0049] Furthermore, the specific analysis steps of the random parameter Logit model considering mean heterogeneity are as follows:

[0050] (21) The Monte Carlo method was used to select slight delays as the reference delay level for establishing the random parameter Logit model. The multinomial logit model (MNL) was first used, 0.01 was selected as the critical significance level of the parameter, and the stepwise backward method was used to eliminate insignificant variables one by one to complete the model calibration.

[0051] Specifically, the MNL model is a traditional Logit model, which assumes that random items are independent of each other and obey extreme value type I distribution. The utility function is:

[0052] S ni =β ni X n +ε ni

[0053] In the formula, X n Represents the influencing factor vector of the delay data in the nth sample; β ni is the estimated parameter vector; ε ni is the error term;

[0054] From the above formula, the density function formula of the MNL model is:

[0055]

[0056] Since the traditional Logit model does not take into account the effect of various factors on different degrees of traffic delay, there is heterogeneity in different data samples. ni Add a random term Γ i v ni ; At this time, the estimated parameter vector is expressed as:

[0057] β ni =β i +Γ i v ni

[0058] In the formula, β i is a fixed parameter, Γ i represents the coefficient matrix, i.e., the covariance and potential correlation of random parameters, v ni represents a random term with zero mean and identity covariance matrix; the estimated parameter vector β ni is represented as a linear combination of fixed parameters and random terms;

[0059] The multinomial logit model further evolved into the random parameter logit model (RPLHM);

[0060] (22) Substitute significant variables into the RPL model to identify random parameters and estimate them; select the Halton sampling method for the sample, assume that all variables are random parameters, and calibrate the parameter distribution in the form of normal distribution and binomial distribution; retain the significant random parameter variables, and set those with significance higher than the threshold as fixed parameters;

[0061] (23) Given the significant parameters of the RPL model, calibrate the random parameter Logit model (RPLHM); maintain the random parameter state of the heterogeneous variables, assume that all fixed parameters are mean heterogeneous variables, and calibrate the parameters in turn; the variables that pass the significance test are calibrated as explanatory variables of the random parameters, and those that fail the significance test remain as fixed parameters;

[0062] (24) Combined with the average marginal effect index, quantitative analysis was conducted on the impact of the four-dimensional factors of road, weather, POI (point of interest) and time on the probability of occurrence of three types of traffic delays: minor, severe and extreme;

[0063] The average marginal effect refers to the average value of the change in the probability of a certain degree of traffic delay caused by a one-unit change in an influencing factor while keeping the values ​​of other variables unchanged. The calculation formula of the average marginal effect value is different for continuous variables and categorical variables.

[0064] In continuous variables, the average marginal effect is the partial derivative of the probability function with respect to the independent variable, and the marginal benefit is calculated as follows:

[0065]

[0066] In the formula, is the average marginal effect of the mth factor on the delay severity j; N is the sample size, n is the nth sample data; x nj is the probability function with the mth variable as the independent variable; x nm Indicates that the independent variable is the mth variable; P n It is the probability function of the nth sample with the kth variable as the independent variable; are tiny differences or increments used in numerical simulations of partial derivatives;

[0067] In the case of binary categorical variables, the average marginal effect is the change in the predicted probability of the degree of delay when the binary variable changes from 0 to 1. The calculation formula is as follows:

[0068]

[0069] Where P nj (x nm =1) is the mth factor variable x of the nth data sample nm =1, the probability of causing traffic delay severity j; nj (x nm =0) is the mth factor variable x of the nth data sample nm =0, the probability of causing traffic delay severity j.

[0070] Furthermore, the specific analysis steps of the Logit model based on random parameters and transferability theory are as follows:

[0071] (31) All sample data are divided into two groups: daytime / nighttime and weekday / weekend;

[0072] (32) Calculate the first log-likelihood ratio test, with a critical confidence level of 95.00%, to test whether there is transferability of influencing factors between the overall traffic delay data and the sub-groups; the sub-groups include daytime, nighttime, weekdays and weekends.

[0073] The test statistic conforms to the chi-square distribution, and the log-likelihood ratio test formula is as follows:

[0074]

[0075] In the formula, is the chi-square value, LL(β joint ) represents the log-likelihood ratio when the model converges based on the delay data caused by traffic accidents of the entire sample; and They are subgroups c i and subgroup c j Log-likelihood ratios at convergence of modeling traffic delay data;

[0076] (33) The second log-likelihood ratio test was calculated with a critical confidence level of 95.00% to compare the local transferability of influencing factors between day and night, between weekdays and weekends, and to analyze the stability of parameter estimates;

[0077] The test statistic of this group conforms to the chi-square distribution, and the calculation formula is as follows:

[0078]

[0079] In the formula, is the chi-square value, Indicates the user group c i Traffic delay degree model for group c j The log-likelihood function value obtained when the traffic data modeling converges; For user group c i The log-likelihood function value at the time of convergence of the traffic delay data modeling;

[0080] (34) Compare the parameter estimation coefficients, signs, and average marginal effects of the influencing factors of different groups. Assuming that the variables with a current significance probability lower than the significance level of 0.01 are significant variables, they are divided into the following three situations: ① significant in one group and not significant in the other; ② the same variable is significant in both groups, weekdays and weekends or daytime and nighttime, but the parameter estimation coefficient signs and average marginal effects have opposite effects; ③ the same variable is significant in both groups, and the parameter estimation coefficient signs and average marginal effects are the same; conduct a difference study on the weekday / weekend and daytime / nighttime groups.

[0081] Furthermore, the specific analysis steps of the latent class and machine learning two-layer model are as follows:

[0082] (41) Establish a latent class model to identify heterogeneous groups, calculate three types of fitting indices (AIC, BIC, SSA-BIC) and three types of significance indices (entropy, LMRT, BLRT) to determine the optimal number of clusters and divide heterogeneous groups;

[0083] Assuming that there are K latent categories in the latent class Logit model, under category k, the conditional probability of the occurrence of the nth event with a traffic delay severity of j is as follows:

[0084]

[0085] In the formula, ε njk represents the error term, which is a constant; X nj represents the set of influencing factor variables of traffic delay severity j in accident sample n; is the independent variable parameter vector of traffic delay severity j under category k; nj|k represents the nth event of traffic delay level j under category k, T represents the transpose of the parameter vector;

[0086] The probability calculation of the traffic delay sample after the nth traffic accident belonging to the k category is as follows:

[0087]

[0088] In the formula, z n is the set of influencing factors belonging to category k; is the corresponding influencing factor variable;

[0089] Combining the above formula, the total probability calculation formula for the occurrence of the nth accident with a traffic delay severity of j caused by a traffic accident is as follows:

[0090]

[0091] According to the above formula, the maximum likelihood method is used to estimate the parameters of the latent class Logit model, and finally the cluster probability corresponding to each traffic delay sample is obtained, and the cluster categories are automatically divided;

[0092] Among them, the latent class model indicators are:

[0093] 1) Akaike Information Criterion (AIC), which is a weighted function of the log-likelihood function value and the number of parameters, is calculated as follows:

[0094] AIC=-2LL(β)+2E

[0095] In the formula, E is the number of free parameters for model convergence. The smaller the value of the AIC index, the better the fitting effect of the model.

[0096] 2) Bayesian Information Criterion (BIC)

[0097] Consider the number of samples to prevent the model from being too complex and overfitting due to excessive model accuracy;

[0098] The calculation formula is as follows:

[0099] BIC=-2LL(β)+Eln(N)

[0100] 3) The sample size-adjusted BIC (SSA-BIC) calculation formula is as follows:

[0101] SSA-BIC=tln((N+2) / 24)-2ln LL(β)

[0102] Where, t is the number of model parameter estimates; LL(β) is the log-likelihood value when the latent class model converges; N is the number of samples, that is, the total sample data of traffic delay data; (N+2) / 24 is the sample size of the corrected traffic delay data; the smaller the SSA-BIC value, the better the model fitting effect;

[0103] 4) Entropy

[0104] Entropy is often used to measure the optimal number of clusters. The calculation formula is as follows:

[0105]

[0106] In the formula, Q represents the number of selected clusters, N is the total number of samples of traffic delays caused by traffic accidents, and p k|n is the posterior probability value that the traffic delay data caused by the nth traffic accident belongs to cluster q; the entropy value is between 0 and 1, and the larger the value, the more accurate the cluster number;

[0107] 5) Likelihood Ratio Test (LMRT) and Bootstrap-based Likelihood Ratio Test (BLRT)

[0108] LMRT and BLRT are used to compare the fitting differences between K-1 and K category models. When the significance probability of the two parameters is significant, the fitting effect of the latent category model corresponding to the current K categories is better than that of the model of K-1 categories. Otherwise, the classification result of the latent category model of K-1 categories is more likely to be selected.

[0109] (42) The SmoteEnn algorithm was used to balance the data, and the feature importance output by the XGboost algorithm was arranged in descending order; for the sorted features, the number of features was gradually reduced at intervals of 10% and the model was trained. The optimal feature ratio of different heterogeneous groups was determined based on the accuracy of the XGboost model, and redundant feature variables were screened;

[0110] (43) Based on the screening of the number of features, the degree of traffic delay at the location of urban traffic accidents was used as the dependent variable, and four models, namely random forest, XGBoost, LightGBM and CatBoost, were used to model and predict the heterogeneous population; the Bayesian model based on tree structure probability density and cross-validation were used to adjust the hyperparameters of the above models;

[0111] (44) Finally, the model performance is evaluated by combining the Accuracy and F1 Score indicators to select the model suitable for each heterogeneous group. The indicators include accuracy, precision and recall:

[0112] (45) Combined with the SHAP method, the SHAP absolute values ​​of different features were calculated to analyze their global importance of features in different groups and the impact of single and interactive features in different heterogeneous groups;

[0113] SHAP is a framework based on game theory to explain the input-output mechanism of machine learning models, which can be combined with different machine learning models for interpretation. SHAP will calculate the distribution attribution value (Shap Value) of the feature for the prediction result of a single sample, and further use the sum of the Shaply values ​​of each feature to explain the model prediction value.

[0114] Specifically:

[0115] First, we need to calculate the Shaply value. Considering the non-independence between features, SHAP will weight the average of possible feature combinations to obtain the final attribution value of feature c. The calculation method is shown in the formula:

[0116]

[0117] In the formula, F is the set of all features, S is a subset of set F and does not include feature c; by training two models, the SHAP value of feature c is extracted, f s∪{c} (x s∪{c} ) represents the model trained using a subset containing feature c, f S (x S ) represents the model trained on feature subset S, and f S (x S )=E[f(x s )|x S], is the expected value of the function of the input feature;

[0118] Secondly, based on the excellent properties of SHAP local accuracy, missingness and consistency, the calculation formula for explaining the predicted value h(x) of the machine learning model with the highest prediction accuracy is obtained, namely:

[0119]

[0120] Where g() is the explanation model, x'∈{0,1} T is the state of whether the feature is observed, T represents the dimension of x', that is, the size of the feature space. If the feature is not observed, then M is the total number of features, f0 is the mean prediction of all samples, and f c is the feature x' c The attributed value of .

[0121] Further, in step (43), the tree structure probability density based Bayesian hyperparameter optimization (TPE-BO) adjustment model includes Bayesian optimization (BO) and tree structure probability density estimation (TPE);

[0122] Bayesian optimization solves the optimal parameter combination by finding the maximum value of the objective function, as shown in the formula:

[0123]

[0124] In the formula, x * To obtain the optimal hyperparameters, c is the parameter space, and f(x) is the objective function;

[0125] Bayesian optimization uses probability distribution model and acquisition function to achieve x * Different Bayesian optimization models are also distinguished by the differences in these two parts. The former is a model established through existing historical parameter adjustment data; the latter provides a reference for the selection of hyperparameter combinations, continuously brings new combinations into the probability distribution model, and iterates to the maximum number of times;

[0126] In the Bayesian optimization model, the acquisition function takes into account the exploration of unknown areas and existing areas, and takes the form of maximizing the expected improvement (EI), as shown in the following formula:

[0127]

[0128] In the formula, y * =min{(x1,f(x1)),2,(x n ,f(x n))}, is the optimal value of the observation threshold; x is the input variable, and y is the predicted input value of the objective function; represents the expected improvement, i.e., the expected improvement value at point x for the target value y, p M (y|x) is the acquisition function;

[0129] Tree structured probability density estimation (TPE) is a probabilistic proxy model. The main difference is that the acquisition function p M (y|x) solution; According to Bayes' theorem, the calculation of this conditional probability can be solved using the formula:

[0130]

[0131] In the formula, x is the input variable, y is the objective function value, p(x) is the prior probability of the input variable x, p(y) is the marginal distribution of the target value y, and p(x|y) is the posterior probability of the input variable x;

[0132] The TPE probabilistic proxy model takes into account the high computational cost of p(y|x) and solves it through p(x|y) and p(y):

[0133]

[0134] p(y <y * )=γ

[0135] Where l(x) is the observed value x i The expected loss function is greater than y * Density estimation of ;

[0136] j(x) is the density estimate of the remaining observations; γ is the y selected in the TPE algorithm * As the quantile value of some observations y;

[0137] From the above formula, the final maximum expected improvement EI can be obtained as:

[0138]

[0139] From the above formula, we can see that the TPE-BO model does not need to model p(y), and the optimization process only Evaluate and obtain the optimal hyperparameter combination.

[0140] The present invention has the following beneficial effects:

[0141] Based on the traffic delay data of urban traffic accident sites, the present invention extracts the road dimension characteristics when the accident occurs, combines the characteristics of missing features, proposes a targeted data filling scheme, completes the redundant feature screening according to the statistical distribution, and constructs structured traffic modeling data with four dimensional characteristics of weather, road, POI and time; systematically models and analyzes the degree of traffic delay caused by urban traffic accidents from three aspects: individual heterogeneity, group transferability and group heterogeneity, and explores the influence of significant variables in multi-dimensional influencing factors, in order to help traffic management departments take preventive measures and emergency response countermeasures, and provide a theoretical basis for road design optimization and management decisions, thereby alleviating the impact of traffic accidents, improving the operating risk level of accident sections, avoiding secondary accidents, and reducing traffic pollution. BRIEF DESCRIPTION OF THE DRAWINGS

[0142] Figure 1 It is a schematic diagram of the analysis of the present invention. DETAILED DESCRIPTION

[0143] The specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings. It should be pointed out that the embodiments are only specific explanations of the invention and should not be regarded as limitations of the invention. The purpose of the embodiments is to enable those skilled in the art to better understand and reproduce the technical solutions of the present invention. The protection scope of the present invention shall still be based on the scope defined by the claims.

[0144] Based on the traffic delay data of urban traffic accident sites, the present invention extracts the road dimension features at the time of the accident, combines the characteristics of missing features, proposes a targeted data filling scheme, completes redundant feature screening according to statistical distribution, and constructs structured traffic modeling data with four dimensional features of weather, road, POI, and time.

[0145] like Figure 1 As shown, the present invention provides a multi-dimensional characteristic analysis system for the degree of traffic delay at a location of an urban traffic accident, and the analysis system includes:

[0146] A preprocessing module is used to preprocess the traffic delay degree dataset of urban accident sites;

[0147] The specific analysis steps are:

[0148] (11) Based on the delay duration, the dependent variable, the degree of traffic delay, is divided into minor delay, severe delay, and extreme delay. Factors that have no explanatory significance and belong to the accident lag variables are preliminarily excluded, including variables such as the accident end time and the starting and ending longitude and latitude.

[0149] (12) Based on the accident description field and the urban road location, the road status, road type and road section characteristic information at the time of the accident are extracted to obtain derived variables including four categories: road status, road type, peak time and season;

[0150] (13) According to the characteristics of statistical distribution, the wind chill temperature and rainfall variable features with too high data loss rate and similar meanings to other variables were deleted; the variance values ​​of all variables were calculated, and the variables with a variance of 0 were deleted, that is, the variables with a single distribution and no discrimination, including speed bumps, roundabouts, turning loops, traffic lights and speed reduction; weather conditions, wind direction, air pressure, visibility, relative humidity, humidity and wind speed were filled;

[0151] The method for filling in weather conditions, wind direction, air pressure, visibility, relative humidity, humidity and wind speed is:

[0152] Cold platform data and mode were used to fill missing values ​​for weather conditions and wind direction variables;

[0153] The Missforest multiple interpolation filling method based on the random forest model is used to perform regression prediction filling for the air pressure, visibility, temperature, relative humidity and wind speed variables. The specific steps of filling are:

[0154] First, combine P variables with n data to be filled into X = (X1, X2, X3, ..., X p ) of n*P dimensional data, use the mode or mean for preliminary filling;

[0155] Secondly, sort the variables with missing rates according to their missing rates and select the variable X with the smallest missing rate. S As the primary fill variable;

[0156] Furthermore, with variable X S As the predictor variable, the remaining variables are used as feature variables and are divided into X according to whether they contain missing values. S Observed value X S Missing values X S Observations other than With X S The remaining observations other than the missing values

[0157] Finally, and Bring it into the random forest model for training, using Data to predict missing data

[0158] During multiple iterations of the random forest model, when the latest filling result changes from the previous filling result by X S When it is infinitely close to 0, the calculation stops and the missing value filling is completed.

[0159]

[0160] In the formula, and Respectively represent X S The latest and last filled data sets of the variable.

[0161] (14) Discretize and bin the temperature, visibility, and wind speed variables, and convert all subcategories of weather conditions, wind direction, peak hours, seasons, temperature, visibility, and wind speed into independent binary features, using 0 and 1 to represent binary state variables;

[0162] The subcategories of weather conditions, wind direction, peak hours, seasons, temperature, visibility, and wind speed are all converted into independent binary features as follows:

[0163] A multi-class variable is divided into Y1, Y2, ...Y i ...,Y d When there are d separate sub-categorical variables, a sub-category Y d It is set as a reference variable and will not be included in the modeling variable range. Then, the binary variable state is represented by 0 and 1. When the variable value is 1, it means that the variable state exists. In addition, in the multi-classification variable, if a subclass Y d When Y is a reference variable, when the values ​​of other subclasses are 0, d The state indicated is established.

[0164] (15) Based on the tolerance and variance inflation factor indicators of the regression model, collinearity diagnosis is performed on all variables in the road dimension, point of interest (POI) dimension, weather dimension, and time dimension to avoid the impact of excessive collinearity levels on the significance level and convergence of the model. The evaluation indicators are tolerance and variance inflation factor (VIF).

[0165] The tolerance is calculated as follows:

[0166]

[0167] In the formula, To make the X i The variables of delay degree are taken as dependent variables, and the remaining k-1 variables are taken as independent variables to establish the determination coefficient of linear regression model; the tolerance value range is between 0 and 1, and the closer it is to 1, the lower the degree of collinearity between independent variables; it is generally believed that when the tolerance is greater than 0.1, there is no multicollinearity between variables, and this paper also selects 0.1 as the indicator.

[0168] The variance inflation factor (VIF) can reflect whether there is multicollinearity between the observed values ​​of the independent variables and is calculated as follows:

[0169]

[0170] In the formula, VIF is the inverse of tolerance, and variable X i of and There is a positive correlation, The larger the value is, the stronger the correlation between the variable and other variables is. The criteria for judging collinearity are:

[0171]

[0172]

[0173] When standard ① is met, it means that there is multicollinearity among the variables; when standard ① and standard ② are met at the same time, it means that the data has serious multicollinearity.

[0174] Based on the random parameter logit model (RPLHM) considering mean heterogeneity, it is used to analyze the individual heterogeneity of the degree of traffic delay at the location of urban traffic accidents; that is, the differences or instability of the influencing factors in different samples, the differences in some characteristics of the samples lead to the different effects of the same influencing factor on different samples, and find out the individual heterogeneous variables and some sample characteristics that cause individual heterogeneity; influencing factor analysis, that is, combined with the average marginal effect value, analyze the variables that have a significant impact on traffic delays from four dimensions: road, weather, points of interest and time, and their differences in impact;

[0175] The specific analysis steps are:

[0176] (21) The Monte Carlo method was used to select slight delays as the reference delay level for establishing the random parameter Logit model. The multinomial logit model (MNL) was first used, 0.01 was selected as the critical significance level of the parameter, and the stepwise backward method was used to eliminate insignificant variables one by one to complete the model calibration.

[0177] Specifically, the MNL model is a traditional Logit model, which assumes that random items are independent of each other and obey extreme value type I distribution. The utility function is:

[0178] S ni =β ni X n +ε ni

[0179] Where, X n represents the influencing factor vector of the delay data in the nth sample; β ni is the estimated parameter vector; εni is the error term;

[0180] From the above formula, the density function formula of the MNL model is:

[0181]

[0182] Since the traditional Logit model does not take into account the effect of various factors on different degrees of traffic delay, there is heterogeneity in different data samples. ni Add a random term Γ i v ni ; At this time, the estimated parameter vector is expressed as:

[0183] β ni =β i +Γ i v ni

[0184] In the formula, β i is a fixed parameter, Γ i represents the coefficient matrix, i.e., the covariance and potential correlation of random parameters, v ni represents a random term with zero mean and identity covariance matrix; the estimated parameter vector β ni is represented as a linear combination of fixed parameters and random terms;

[0185] The multinomial logit model further evolved into the random parameter logit model (RPLHM);

[0186] (22) Substitute significant variables into the RPL model to identify random parameters and estimate them; select the Halton sampling method for the sample, assume that all variables are random parameters, and calibrate the parameter distribution in the form of normal distribution and binomial distribution; retain the significant random parameter variables, and set those with significance higher than the threshold as fixed parameters;

[0187] (23) Given the significant parameters of the RPL model, calibrate the random parameter Logit model (RPLHM); maintain the random parameter state of the heterogeneous variables, assume that all fixed parameters are mean heterogeneous variables, and calibrate the parameters in turn; the variables that pass the significance test are calibrated as explanatory variables of the random parameters, and those that fail the significance test remain as fixed parameters;

[0188] (24) Combined with the average marginal effect index, quantitative analysis was conducted on the impact of the four-dimensional factors of road, weather, POI (point of interest) and time on the probability of occurrence of three types of traffic delays: minor, severe and extreme;

[0189] The average marginal effect refers to the average value of the change in the probability of a certain degree of traffic delay caused by a one-unit change in an influencing factor while keeping the values ​​of other variables unchanged. The calculation formula of the average marginal effect value is different for continuous variables and categorical variables.

[0190] In continuous variables, the average marginal effect is the partial derivative of the probability function with respect to the independent variable, and the marginal benefit is calculated as follows:

[0191]

[0192] In the formula, is the average marginal effect of the mth factor on the delay severity j; N is the sample size, n is the nth sample data; x nj is the probability function with the mth variable as the independent variable; x nm Indicates that the independent variable is the mth variable; P n It is the probability function of the nth sample with the kth variable as the independent variable; are tiny differences or increments used in numerical simulations of partial derivatives;

[0193] In the case of binary categorical variables, the average marginal effect is the change in the predicted probability of the degree of delay when the binary variable changes from 0 to 1. The calculation formula is as follows:

[0194]

[0195] Where P nj (x nm =1) is the mth factor variable x of the nth data sample nm =1, the probability of causing traffic delay severity j; nj (x nm =0) is the mth factor variable x of the nth data sample nm =0, the probability of causing traffic delay severity j.

[0196] Based on the random parameter logit model (RPLHM), the overall and local transferability analysis of artificially divided groups is carried out by combining with the transferability theory, that is, whether the influencing mechanisms of the two groups are similar, and whether the model of a single group can maintain good performance on the data of another group; the overall data set is divided into day / night and weekdays / holidays, and the consistency or difference of the influencing factors of traffic delays between groups is detected.

[0197] The specific analysis steps are:

[0198] (31) All sample data are divided into two groups: daytime / nighttime and weekday / weekend;

[0199] (32) Calculate the first log-likelihood ratio test, with a critical confidence level of 95.00%, to test whether there is transferability of influencing factors between the overall traffic delay data and the sub-groups; the sub-groups include daytime, nighttime, weekdays and weekends.

[0200] The test statistic conforms to the chi-square distribution, and the log-likelihood ratio test formula is as follows:

[0201]

[0202] In the formula, is the chi-square value, LL(β joint ) represents the log-likelihood ratio when the model converges based on the delay data caused by traffic accidents of the entire sample; and They are subgroups c i and subgroup c j Log-likelihood ratios at convergence of modeling traffic delay data;

[0203] (33) The second log-likelihood ratio test was calculated with a critical confidence level of 95.00% to compare the local transferability of influencing factors between day and night, between weekdays and weekends, and to analyze the stability of parameter estimates;

[0204] The test statistic of this group conforms to the chi-square distribution, and the calculation formula is as follows:

[0205]

[0206] In the formula, is the chi-square value, Indicates the user group c i Traffic delay degree model for group c j The log-likelihood function value obtained when the traffic data modeling converges; For user group c i The log-likelihood function value at the time of convergence of the traffic delay data modeling;

[0207] (34) Compare the parameter estimation coefficients, signs, and average marginal effects of the influencing factors of different groups. Assuming that the variables with a current significance probability lower than the significance level of 0.01 are significant variables, they are divided into the following three situations: ① significant in one group and not significant in the other; ② the same variable is significant in both groups, weekdays and weekends or daytime and nighttime, but the parameter estimation coefficient signs and average marginal effects have opposite effects; ③ the same variable is significant in both groups, and the parameter estimation coefficient signs and average marginal effects are the same; conduct a difference study on the weekday / weekend and daytime / nighttime groups.

[0208] The latent class and machine learning double-layer model is used to identify heterogeneous groups of traffic delays in urban traffic accident sites and to clarify the best scenario for dividing heterogeneous groups; delay degree prediction, select the best machine learning model to predict the delay degree of heterogeneous groups; group heterogeneity analysis, combined with the SHAP model to explain and analyze the global importance of different features in heterogeneous groups.

[0209] The specific analysis steps are:

[0210] (41) Establish a latent class model to identify heterogeneous groups, calculate three types of fitting indices (AIC, BIC, SSA-BIC) and three types of significance indices (entropy, LMRT, BLRT) to determine the optimal number of clusters and divide heterogeneous groups;

[0211] Assuming that there are K latent categories in the latent class Logit model, under category k, the conditional probability of the occurrence of the nth event with a traffic delay severity of j is as follows:

[0212]

[0213] In the formula, ε njk represents the error term, which is a constant; X nj represents the set of influencing factor variables of traffic delay severity j in accident sample n; is the independent variable parameter vector of traffic delay severity j under category k; njk represents the nth event of traffic delay level j under category k, T represents the transpose of the parameter vector;

[0214] The probability calculation of the traffic delay sample after the nth traffic accident belonging to the k category is as follows:

[0215]

[0216] In the formula, z n is the set of influencing factors belonging to category k; is the corresponding influencing factor variable;

[0217] Combining the above formula, the total probability calculation formula for the occurrence of the nth accident with a traffic delay severity of j caused by a traffic accident is as follows:

[0218]

[0219] According to the above formula, the maximum likelihood method is used to estimate the parameters of the latent class Logit model, and finally the cluster probability corresponding to each traffic delay sample is obtained, and the cluster categories are automatically divided;

[0220] Among them, the latent class model indicators are:

[0221] 1) Akaike Information Criterion (AIC), which is a weighted function of the log-likelihood function value and the number of parameters, is calculated as follows:

[0222] AIC=-2LL(β)+2E

[0223] In the formula, E is the number of free parameters for model convergence. The smaller the value of the AIC index, the better the fitting effect of the model.

[0224] 2) Bayesian Information Criterion (BIC)

[0225] Consider the number of samples to prevent the model from being too complex and overfitting due to excessive model accuracy;

[0226] The calculation formula is as follows:

[0227] BIC=-2LL(β)+Eln(N)

[0228] 3) The sample size-adjusted BIC (SSA-BIC) calculation formula is as follows:

[0229] SSA-BIC=tln((N+2) / 24)-2ln LL(β)

[0230] Where, t is the number of model parameter estimates; LL(β) is the log-likelihood value when the latent class model converges; N is the number of samples, that is, the total sample data of traffic delay data; (N+2) / 24 is the sample size of the corrected traffic delay data; the smaller the SSA-BIC value, the better the model fitting effect;

[0231] 4) Entropy

[0232] Entropy is often used to measure the optimal number of clusters. The calculation formula is as follows:

[0233]

[0234] In the formula, Q represents the number of selected clusters, N is the total number of samples of traffic delays caused by traffic accidents, and p k|n is the posterior probability value that the traffic delay data caused by the nth traffic accident belongs to cluster q; the entropy value is between 0 and 1, and the larger the value, the more accurate the cluster number;

[0235] 5) Likelihood Ratio Test (LMRT) and Bootstrap-based Likelihood Ratio Test (BLRT)

[0236] LMRT and BLRT are used to compare the fitting differences between K-1 and K category models. When the significance probability of the two parameters is significant, the fitting effect of the latent category model corresponding to the current K categories is better than that of the model of K-1 categories. Otherwise, the classification result of the latent category model of K-1 categories is more likely to be selected.

[0237] (42) The SmoteEnn algorithm was used to balance the data, and the feature importance output by the XGboost algorithm was arranged in descending order; for the sorted features, the number of features was gradually reduced at intervals of 10% and the model was trained. The optimal feature ratio of different heterogeneous groups was determined based on the accuracy of the XGboost model, and redundant feature variables were screened;

[0238] (43) Based on the screening of the number of features, the degree of traffic delay at the location of urban traffic accidents was used as the dependent variable, and four models, namely random forest, XGBoost, LightGBM and CatBoost, were used to model and predict the heterogeneous population; the Bayesian model based on tree structure probability density and cross-validation were used to adjust the hyperparameters of the above models;

[0239] The tree-structured probability density-based Bayesian hyperparameter optimization (TPE-BO) adjustment model includes Bayesian Optimization (BO) and Tree Parzer Estimator (TPE);

[0240] Bayesian optimization solves the optimal parameter combination by finding the maximum value of the objective function, as shown in the formula:

[0241]

[0242] In the formula, x * To obtain the optimal hyperparameters, c is the parameter space, and f(x) is the objective function;

[0243] Bayesian optimization uses probability distribution model and acquisition function to achieve x * Different Bayesian optimization models are also distinguished by the differences in these two parts. The former is a model established through existing historical parameter adjustment data; the latter provides a reference for the selection of hyperparameter combinations, continuously brings new combinations into the probability distribution model, and iterates to the maximum number of times;

[0244] In the Bayesian optimization model, the acquisition function takes into account the exploration of unknown areas and existing areas, and takes the form of maximizing the expected improvement (EI), as shown in the following formula:

[0245]

[0246] In the formula, y * =min{(x1,f(x1)),2,(x n ,f(x n ))}, is the optimal value of the observation threshold; x is the input variable, and y is the predicted input value of the objective function; represents the expected improvement, i.e., the expected improvement value at point x for the target value y, p M (y|x) is the acquisition function;

[0247] Tree structured probability density estimation (TPE) is a probabilistic proxy model. The main difference is that the acquisition function p M (y|x) solution; According to Bayes' theorem, the calculation of this conditional probability can be solved using the formula:

[0248]

[0249] In the formula, x is the input variable, y is the objective function value, p(x) is the prior probability of the input variable x, p(y) is the marginal distribution of the target value y, and p(x|y) is the posterior probability of the input variable x;

[0250] The TPE probabilistic proxy model takes into account the high computational cost of p(y|x) and solves it through p(x|y) and p(y):

[0251]

[0252] p(y <y * )=γ

[0253] Where l(x) is the observed value x i The expected loss function is greater than y * Density estimation of ;

[0254] j(x) is the density estimate of the remaining observations; γ is the y selected in the TPE algorithm * As the quantile value of some observations y;

[0255] From the above formula, the final maximum expected improvement EI can be obtained as:

[0256]

[0257] From the above formula, we can see that the TPE-BO model does not need to model p(y), and the optimization process only Evaluate and obtain the optimal hyperparameter combination.

[0258] (44) Finally, the model performance is evaluated by combining the Accuracy and F1 Score indicators to select the model suitable for each heterogeneous group. The indicators include accuracy, precision and recall:

[0259] (45) Combined with the SHAP method, the SHAP absolute values ​​of different features were calculated to analyze their global importance of features in different groups and the impact of single and interactive features in different heterogeneous groups;

[0260] SHAP is a framework based on game theory to explain the input-output mechanism of machine learning models, which can be combined with different machine learning models for interpretation. SHAP will calculate the distribution attribution value (Shap Value) of the feature for the prediction result of a single sample, and further use the sum of the Shaply values ​​of each feature to explain the model prediction value.

[0261] Specifically:

[0262] First, we need to calculate the Shaply value. Considering the non-independence between features, SHAP will weight the average of possible feature combinations to obtain the final attribution value of feature c. The calculation method is shown in the formula:

[0263]

[0264] In the formula, F is the set of all features, S is a subset of set F and does not include feature c; by training two models, the SHAP value of feature c is extracted, f s∪{c} (x s∪{c} ) represents the model trained using a subset containing feature c, f S (x S ) represents the model trained on feature subset S, and f S (x S )=E[f(x s )x S ], is the expected value of the function of the input feature;

[0265] Secondly, based on the excellent properties of SHAP local accuracy, missingness and consistency, the calculation formula for explaining the predicted value h(x) of the machine learning model with the highest prediction accuracy is obtained, namely:

[0266]

[0267] Where g() is the explanation model, x'∈{0,1} T is the state of whether the feature is observed, T represents the dimension of x', that is, the size of the feature space. If the feature is not observed, then M is the total number of features, f0 is the mean prediction of all samples, and f c is the feature x' c The attributed value of .

[0268] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0269] It should be noted that the technical features not described in detail in the present invention can be implemented by any existing technology.

Claims

1. An analysis system for multi-dimensional characteristics of traffic delay degree at urban traffic accident sites, the analysis system comprising: A preprocessing module is used to preprocess the traffic delay degree dataset of urban accident sites; The random parameter logit model considering mean heterogeneity is used to analyze the individual heterogeneity of the degree of traffic delay at urban traffic accident sites. Based on the Logit model with random parameters, the overall and local transferability analysis of artificially divided groups is conducted by combining it with the transferability theory. The latent class and machine learning two-layer model is used to identify heterogeneous groups of traffic delays in urban traffic accident sites and to identify the best scenario for dividing heterogeneous groups.

2. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 1 is characterized in that: The specific analysis steps of the preprocessing model are: (11) According to the delay duration, the dependent variable, the degree of traffic delay, is divided into minor delay, severe delay, and extreme delay. Factors that have no explanatory significance and belong to the accident lag variables are preliminarily excluded, including the accident end time and the start and end latitude and longitude variables; (12) Based on the accident description field and the urban road location, four types of derived variables are extracted: road status, road type, peak hour, and season at the time of the accident; (13) According to the statistical distribution characteristics, the wind chill temperature and rainfall variable features with too high data missing rate and similar meanings to other variables were deleted; Calculate the variance values ​​of all variables and delete variables with a variance of 0, that is, variables with a single distribution and no discrimination, including speed bumps, yield signs, roundabouts, turning loops, traffic lights, and speed reduction; fill in the variables of weather conditions, wind direction, air pressure, visibility, relative humidity, humidity, and wind speed; (14) Discretize and bin the temperature, visibility, and wind speed variables, and convert all subcategories of weather conditions, wind direction, peak hours, seasons, temperature, visibility, and wind speed into independent binary features, using 0 and 1 to represent binary state variables; (15) Based on the tolerance and variance inflation factor indicators of the regression model, collinearity diagnosis is performed on all variables in the road dimension, point of interest, dimension, weather dimension, and time dimension.

3. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 2 is characterized in that: In step (13), the method for filling in the weather conditions, wind direction, air pressure, visibility, relative humidity, humidity and wind speed is: Cold platform data and mode were used to fill missing values ​​for weather conditions and wind direction variables; The Missforest multiple interpolation filling method based on the random forest model is used to perform regression prediction filling for the air pressure, visibility, temperature, relative humidity and wind speed variables. The specific steps of filling are: First, combine P variables with n data to be filled into X = (X1, X2, X3, ..., X p ) of n*P dimensional data, use the mode or mean for preliminary filling; Secondly, sort the variables with missing rates according to their missing rates and select the variable X with the smallest missing rate. S As the primary fill variable; Furthermore, with variable X S As the predictor variable, the remaining variables are used as feature variables and are divided into X according to whether they contain missing values. S Observed value X S Missing values X S Observations other than With X S The remaining observations other than the missing values Finally, and Bring it into the random forest model for training, using Data to predict missing data During multiple iterations of the random forest model, when the latest filling result changes from the previous filling result by X S When it is infinitely close to 0, the calculation stops and the missing value filling is completed. In the formula, and Respectively represent X S The latest and last filled data sets of the variable.

4. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 2 is characterized in that: In step (14), all subcategories of weather conditions, wind direction, peak hours, seasons, temperature, visibility, and wind speed are converted into independent binary features by: A multi-class variable is divided into Y1, Y2, ...Y i ...,Y d When there are d separate sub-categorical variables, a sub-category Y d It is set as a reference variable and will not be included in the modeling variable range. Then, the binary variable state is represented by 0, 1. When the variable value is 1, it means that the variable state exists; In addition, in a multi-class variable, if a subclass Y d When Y is a reference variable, when the values ​​of other subclasses are 0, d The state indicated is established.

5. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 2 is characterized in that: In step (15), the evaluation indicators of collinearity diagnosis are tolerance and variance inflation factor; The tolerance is calculated as follows: In the formula, To make the X i The variables of the degree of delay are taken as dependent variables, and the remaining k-1 variables are taken as independent variables to establish the determination coefficient of the linear regression model; the tolerance value range is between 0 and 1, and the closer it is to 1, the lower the degree of collinearity between the independent variables; The variance inflation factor is calculated as follows: In the formula, VIF is the inverse of tolerance, and variable X i of and There is a positive correlation, The larger it is, the stronger the correlation between the variable and other variables; The criteria for judging collinearity are: ① ② When standard ① is met, it means that there is multicollinearity among the variables; when standard ① and standard ② are met at the same time, it means that the data has serious multicollinearity.

6. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 1, characterized in that: The specific analysis steps based on the random parameter Logit model considering mean heterogeneity are: (21) The Monte Carlo method was used to select slight delays as the reference delay level for establishing the random parameter Logit model. The multinomial logit model was first used, 0.01 was selected as the critical significance level of the parameter, and the stepwise backward method was used to eliminate insignificant variables one by one to complete the model calibration. Specifically, the MNL model is a traditional Logit model, which assumes that random items are independent of each other and obey extreme value type I distribution. The utility function is: S ni =b ni X n +e ni Where, X n Represents the influencing factor vector of the delay data in the nth sample; β ni is the estimated parameter vector; ε ni is the error term; From the above formula, the density function formula of the MNL model is: Since the traditional Logit model does not take into account the effect of various factors on different degrees of traffic delay, there is heterogeneity in different data samples. ni Add a random term Γ i v ni ; At this time, the estimated parameter vector is expressed as: b ni =b i +C i v ni In the formula, β i is a fixed parameter, Γ i represents the coefficient matrix, i.e., the covariance and potential correlation of random parameters, v ni represents a random term with zero mean and identity covariance matrix; the estimated parameter vector β ni is represented as a linear combination of fixed parameters and random terms; The multinomial logit model further evolved into the random parameter logit model; (22) Substitute significant variables into the RPL model to identify random parameters and estimate them; select the Halton sampling method for the sample, assume that all variables are random parameters, and calibrate the parameter distribution in the form of normal distribution and binomial distribution; retain the significant random parameter variables, and set those with significance higher than the threshold as fixed parameters; (23) Given the significant parameters of the RPL model, calibrate the random parameter Logit model; maintain the random parameter state of the heterogeneous variables, assume that all fixed parameters are mean heterogeneous variables, and calibrate the parameters in turn; variables that pass the significance test are calibrated as explanatory variables of the random parameters, and those that fail the significance test remain fixed parameters; (24) Combined with the average marginal effect index, quantitative analysis was conducted on the impact of the four-dimensional factors of road, weather, points of interest, and time on the probability of occurrence of three types of traffic delays: minor, severe, and extreme; In continuous variables, the average marginal effect is the partial derivative of the probability function with respect to the independent variable, and the marginal benefit is calculated as follows: In the formula, is the average marginal effect of the mth factor on the delay severity j; N is the sample size, n is the nth sample data; x nj is the probability function with the mth variable as the independent variable; x nm Indicates that the independent variable is the mth variable; P n It is the probability function of the nth sample with the kth variable as the independent variable; are tiny differences or increments used in numerical simulations of partial derivatives; In the case of binary categorical variables, the average marginal effect is the change in the predicted probability of the degree of delay when the binary variable changes from 0 to 1. The calculation formula is as follows: Where P nj (x nm =1) is the mth factor variable x of the nth data sample nm =1, the probability of causing traffic delay severity j; nj (x nm =0) is the mth factor variable x of the nth data sample nm =0, the probability of causing traffic delay severity j.

7. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 1, characterized in that: The specific analysis steps of the Logit model based on random parameters and transferability theory are as follows: (31) All sample data are divided into two groups: daytime / nighttime and weekday / weekend; (32) Calculate the first log-likelihood ratio test, with a critical confidence level of 95.00%, to test whether there is transferability of influencing factors between the overall traffic delay data and the sub-groups; the sub-groups include daytime, nighttime, weekdays and weekends. The test statistic conforms to the chi-square distribution, and the log-likelihood ratio test formula is as follows: In the formula, is the chi-square value, LL(β joint ) represents the log-likelihood ratio when the model converges based on the delay data caused by traffic accidents of the entire sample; and They are subclass groups c i and subgroup c j Log-likelihood ratios at convergence of modeling traffic delay data; (33) The second log-likelihood ratio test was calculated with a critical confidence level of 95.00% to compare the local transferability of influencing factors between day and night, between weekdays and weekends, and to analyze the stability of parameter estimates; The test statistic of this group conforms to the chi-square distribution, and the calculation formula is as follows: In the formula, is the chi-square value, Indicates the user group c i Traffic delay degree model for group c j The log-likelihood function value obtained when the traffic data modeling converges; For user group c i The log-likelihood function value at the time of convergence of the traffic delay data modeling; (34) Compare the parameter estimation coefficients, signs, and average marginal effects of the influencing factors of different groups. Assuming that the variables with a current significance probability lower than the significance level of 0.01 are significant variables, they are divided into the following three situations: ① significant in one group and not significant in the other; ② the same variable is significant in both groups, weekdays and weekends or daytime and nighttime, but the parameter estimation coefficient signs and average marginal effects have opposite effects; ③ the same variable is significant in both groups, and the parameter estimation coefficient signs and average marginal effects are the same; conduct a difference study on the weekday / weekend and daytime / nighttime groups.

8. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 1, characterized in that: The specific analysis steps of the latent class and machine learning two-layer model are: (41) Establish a latent class model to identify heterogeneous groups, calculate three types of fitting indices (AIC, BIC, SSA-BIC) and three types of significance indices (entropy, LMRT, BLRT) to determine the optimal number of clusters and divide heterogeneous groups; Assuming that there are K latent categories in the latent class Logit model, under category k, the conditional probability of the occurrence of the nth event with a traffic delay severity of j is as follows: In the formula, ε nj|k represents the error term, which is a constant; X nj represents the set of influencing factor variables of traffic delay severity j in accident sample n; is the independent variable parameter vector of traffic delay severity j under category k; nj|k represents the nth event of traffic delay level j under category k, T represents the transpose of the parameter vector; The probability calculation of the traffic delay sample after the nth traffic accident belonging to the k category is as follows: In the formula, z n is the set of influencing factors belonging to category k; is the corresponding influencing factor variable; Combining the above formula, the total probability calculation formula for the occurrence of the nth accident with a traffic delay severity of j caused by a traffic accident is as follows: According to the above formula, the maximum likelihood method is used to estimate the parameters of the latent class Logit model, and finally the cluster probability corresponding to each traffic delay sample is obtained, and the cluster categories are automatically divided; (42) The SmoteEnn algorithm was used to balance the data, and the feature importance output by the XGboost algorithm was arranged in descending order; for the sorted features, the number of features was gradually reduced at intervals of 10% and the model was trained. The optimal feature ratio of different heterogeneous groups was determined based on the accuracy of the XGboost model, and redundant feature variables were screened; (43) Based on the screening of the number of features, the degree of traffic delay at the location of urban traffic accidents was used as the dependent variable, and four models, namely random forest, XGBoost, LightGBM and CatBoost, were used to model and predict the heterogeneous population; the Bayesian model based on tree structure probability density and cross-validation were used to adjust the hyperparameters of the above models; (44) Finally, the model performance is evaluated by combining the Accuracy and F1 Score indicators to select the model suitable for each heterogeneous group. The indicators include accuracy, precision and recall: (45) Combined with the SHAP method, the SHAP absolute values ​​of different features were calculated to analyze their global importance of features in different groups and the impact of single and interactive features in different heterogeneous groups; Specifically: First, we need to calculate the Shaply value. Considering the non-independence between features, SHAP will weight the average of possible feature combinations to obtain the final attribution value of feature c. The calculation method is shown in the formula: In the formula, F is the set of all features, S is a subset of set F and does not include feature c; by training two models, the SHAP value of feature c is extracted, f s∪{c} (x s∪{c} ) represents the model trained using a subset containing feature c, f S (x S ) represents the model trained on feature subset S, and f S (x S )=E[f(x s )|x S ], is the expected value of the function of the input feature; Secondly, based on the excellent properties of SHAP local accuracy, missingness and consistency, the calculation formula for explaining the predicted value h(x) of the machine learning model with the highest prediction accuracy is obtained, namely: Where g() is the explanation model, x'∈{0,1} T is the state of whether the feature is observed, T represents the dimension of x', that is, the size of the feature space. If the feature is not observed, then M is the total number of features, f0 is the mean prediction of all samples, and f c is the feature x' c The attributed value of .

9. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 8, characterized in that: In step (41), the latent class model indicators are: 1) Akaike Information Criterion value is a weighted function of the log-likelihood function value and the number of parameters. The calculation formula is as follows: AIC=-2LL(β)+2E In the formula, E is the number of free parameters for model convergence. The smaller the value of the AIC index, the better the fitting effect of the model. 2) Bayesian Information Criterion Value Consider the number of samples to prevent the model from being too complex and overfitting due to excessive model accuracy; The calculation formula is as follows: BIC=-2LL(β)+Eln(N) 3) The sample-corrected SSA-BIC calculation formula is as follows: SSA-BIC=tln((N+2) / 24)-2lnLL(β) Where t is the number of model parameter estimates; LL(β) is the log-likelihood value when the latent class model converges; N is the number of samples, that is, the total sample data of traffic delay data; (N+2) / 24 is the sample size of the corrected traffic delay data; The smaller the SSA-BIC value, the better the model fit; 4) Entropy Entropy is often used to measure the optimal number of clusters. The calculation formula is as follows: In the formula, Q represents the number of selected clusters, N is the total number of samples of traffic delays caused by traffic accidents, and p k|n is the posterior probability value that the traffic delay data caused by the nth traffic accident belongs to cluster q; the entropy value is between 0 and 1, and the larger the value, the more accurate the cluster number; 5) Likelihood ratio test and Bootstrap-based likelihood ratio test LMRT and BLRT are used to compare the fitting differences between K-1 and K category models. When the significance probabilities of the two parameters are significant, the fitting effect of the latent category model corresponding to the current K categories is better than that of the K-1 category model. Otherwise, the classification results of the latent category model of K-1 categories are more likely to be selected.

10. The method for analyzing the multi-dimensional characteristics of the degree of traffic delay at a location of an urban traffic accident according to claim 8, characterized in that: In step (43), the Bayesian hyperparameter optimization adjustment model based on tree structure probability density includes Bayesian optimization and tree structure probability density estimation; Bayesian optimization solves the optimal parameter combination by finding the maximum value of the objective function, as shown in the formula: In the formula, x * To obtain the optimal hyperparameters, c is the parameter space, and f(x) is the objective function; Bayesian optimization uses probability distribution model and acquisition function to achieve x * Different Bayesian optimization models are also distinguished by the differences in these two parts. The former is a model established through existing historical parameter adjustment data; the latter provides a reference for the selection of hyperparameter combinations, continuously brings new combinations into the probability distribution model, and iterates to the maximum number of times; In the Bayesian optimization model, the acquisition function takes into account the exploration of unknown areas and existing areas, and takes the form of maximizing expected improvement, as shown in the following formula: In the formula, y * =min{(x1,f(x1)),…,(x n ,f(x n ))}, is the optimal value of the observation threshold; x is the input variable, and y is the predicted input value of the objective function; represents the expected improvement, i.e., the expected improvement value at point x for the target value y, p M (y|x) is the acquisition function; Tree structured probability density estimation (TPE) is a probabilistic proxy model. The main difference is that the acquisition function p M (y|x) solution; According to Bayes' theorem, the calculation of this conditional probability can be solved using the formula: In the formula, x is the input variable, y is the objective function value, p(x) is the prior probability of the input variable x, p(y) is the marginal distribution of the target value y, and p(x|y) is the posterior probability of the input variable x; The TPE probabilistic proxy model takes into account the high computational cost of p(y|x) and solves it through p(x|y) and p(y): p(y <y * )=γ Where l(x) is the observed value x i The expected loss function is greater than y * Density estimation of ; j(x) is the density estimate of the remaining observations; γ is the y selected in the TPE algorithm * As the quantile value of some observations y; From the above formula, the final maximum expected improvement EI can be obtained as: From the above formula, we can see that the TPE-BO model does not need to model p(y), and the optimization process only Evaluate and obtain the optimal hyperparameter combination.

Citation Information

Patent Citations

  • Road safety analysis method and system for eliminating traffic accident sample heterogeneity

    CN116384793A

  • Rural highway traffic accident severity evaluation method

    CN117648539A

  • Traffic accident risk studying and judging method based on interpretable random forest

    CN118095834A

  • Multimodal data integration method considering spatiotemporal characteristics of disaster damage

    KR102379472B1

  • Parameterizing Cell-to-Cell Regulatory Heterogeneities via Stochastic Transcriptional Profiles

    US20160253453A1