Traffic accident severity influence factor analysis method based on local cascade integration

Through the local cascade integration method combined with SMOTENC and Hyperopt optimization, the model deviation-variance trade-off problem in the existing traffic accident severity analysis is solved, and higher analysis accuracy and classification performance are achieved, revealing the complex interactive relationship between the accident severity and influencing factors.

CN120277552APending Publication Date: 2025-07-08HARBIN INST OF TECH AT WEIHAI
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510383288.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing traffic accident severity analysis methods are limited by statistical assumptions, resulting in insufficient accuracy and reliability of the analysis results. In addition, machine learning models have bias-variance trade-offs in small sample or high-dimensional data scenarios, making it difficult to accurately reveal the complex interaction between the accident severity and influencing factors.

Method used

The local cascade integration method is adopted, combined with the SMOTENC algorithm and the Hyperopt method, and the deviation-variance trade-off problem is handled through multi-level and multi-stage models, and the model is visualized and interpreted using SHAP tools to analyze the factors affecting the severity of traffic accidents.

Benefits of technology

The model's training data fit and the generalization ability of unknown data are improved. By balancing the training number of various accidents, the model's classification performance and analysis accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277552A_ABST
    Figure CN120277552A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of traffic safety management, and discloses a traffic accident severity influence factor analysis method based on local cascade integration, which comprises the following steps: acquiring an accident data set D1; processing the accident data set D1, removing part of redundant attributes, dividing the accident data set D1 into a training set and a test set according to a proportion, and balancing the number of accidents with different severity degrees in the training set through an SMOTENC algorithm to obtain an accident data set D2; importing the accident data set D2 into a local cascade integration model, adjusting hyper-parameters of a local model through a Hyperpt method, and selecting an optimal hyper-parameter combination by using k-fold cross validation; drawing a confusion matrix according to a training result, and selecting indexes to evaluate model performance; and visualizing the model by applying a machine learning output explanation tool SHAP, and analyzing accident severity influence factors according to the visualized model. By adopting the SMOTENC resampling technology, the number of various accidents in the training set is balanced and the model training effect and the classification performance are improved on the premise of considering discrete and continuous variable differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of traffic safety management, and particularly relates to a method for analyzing influencing factors of traffic accident severity based on local cascade integration. Background Art

[0002] According to the degree of casualties of people in traffic accidents, traffic accidents can be divided into different levels. The lightest accident level is an accident with only property damage (no casualties), and the heaviest accident level is an accident with fatalities. Traffic accidents of different severities have different impacts on economic losses and road traffic. Traffic accidents with a higher severity level not only cause huge social and economic losses, but also seriously affect the operation efficiency of expressways and urban expressways.

[0003] By analyzing the influencing factors of traffic accidents, revealing the influencing mechanism of accident severity, and then taking targeted prevention and control measures is the most direct and effective means to reduce the accident fatality rate, and a reliable analysis method is the premise to ensure the accuracy of the analysis of the influencing mechanism of accident severity.

[0004] Traditional analysis of accident severity mostly uses a discrete choice model based on Logit. Although this method can quantify the influence of various factors on accident severity, it is limited by strict statistical hypothesis conditions (including: ① linear hypothesis of the utility function; ② extreme value distribution hypothesis of the random disturbance term), resulting in great room for improvement in the accuracy and reliability of the analysis results. In contrast, machine learning models do not require any hypothesis conditions and can theoretically reveal the complex interaction relationship between accident severity and influencing factors more accurately, improving the accuracy of accident severity analysis. In the field of traditional machine learning, the bias-variance trade-off has always been one of the core challenges in model design, which reflects the generalization ability of machine learning models. Classic models (such as linear regression) are prone to underfitting due to high bias and ignoring data characteristics, while complex models (such as deep neural networks) are overly sensitive to noise due to high variance and cause overfitting, which is particularly prominent in small sample or high-dimensional data scenarios. Existing methods such as regularization and ensemble learning attempt to alleviate this contradiction through model combination, but they are still limited by the homogeneity of the base learners. Summary of the Invention

[0005] To solve the deficiencies in the prior art, the present invention provides a method for analyzing influencing factors of traffic accident severity based on local cascade integration, which solves the bias-variance trade-off problem through multi-level and multi-stage models and has great potential in accident severity analysis.

[0006] To achieve the above object, the present invention is realized through the following technical solutions:

[0007] A method for analyzing factors affecting the severity of traffic accidents based on local cascade integration includes the following steps:

[0008] S1, obtain the accident data set D1, including driver, vehicle, road and environmental variable information;

[0009] S2, process the accident data set D1, remove some redundant attributes, divide it into a training set and a test set in proportion, and balance the number of accidents of different severity in the training set through the SMOTENC algorithm to obtain the accident data set D2;

[0010] S3, import the accident data set D2 into the local cascade integration model, adjust the hyperparameters of the local model through the Hyperopt method, and use k-fold cross validation to select the best hyperparameter combination;

[0011] S4. Draw a confusion matrix based on the training results and select indicators to evaluate model performance;

[0012] S5. Use the machine learning output interpretation tool SHAP to visualize the model and analyze the factors affecting the severity of the accident.

[0013] Furthermore, the specific method of S2 is:

[0014] S21. Remove accident data that has incomplete information or contains obvious errors;

[0015] S22. Select different correlation coefficients according to the variable type to test their correlation. If the correlation coefficient between two variables is ≥ 0.7, only one of the variables needs to be included in the modeling;

[0016] Among them, Pearson's correlation coefficient p was used to test the correlation between two continuous variables;

[0017] Cramer's correlation coefficient V was used to test the correlation between two discrete variables;

[0018] The Spearman rank correlation coefficient ρ was used to test the correlation between continuous and discrete variables;

[0019] S23. Divide the data set D1 into a training set and a test set in a ratio of 7:3, and use the SMOTENC algorithm to balance the number of accidents of different severity in the training set to obtain the accident data set D2. The specific steps are: first, clarify which discrete variables are in the training set; then, except for the sample with the largest severity (generally "no casualties"), for each of the remaining samples, randomly select a similar sample from the k nearest neighbors, perform linear interpolation for continuous variables, and take the mode of the nearest neighbors for discrete variables; finally, for different types of accidents, synthesize artificial data in a certain proportion.

[0020] Furthermore, the specific method of S3 is:

[0021] S31. The hyperparameters of the model include the maximum tree depth, the number of trees in the ensemble, the number of base learners, the maximum tree depth of the base learners, the learning rate of the base learners, and the minimum loss reduction when splitting nodes.

[0022] S32. Randomly and evenly divide the accident dataset D2 into k = 5 subsets, where 4 subsets are used for training and the remaining 1 subset is used as a validation set to test the model performance. Repeat this process 5 times until the training is completed to find the optimal combination to maximize the model performance.

[0023] Furthermore, the selected evaluation metrics in S4 include Accuracy, Recall, Precision, and F1-score, and the calculation formulas are as follows:

[0024]

[0025]

[0026]

[0027]

[0028] In the formula, TP is the true positive, FN is the false negative, FP is the false positive, and TN is the true negative.

[0029] Furthermore, the specific method of S5 is as follows:

[0030] S51. Select the average value of the training data as the baseline value, that is, the output of the model without any variable information; for the input variable set N = {1, 2,..., M}, calculate all possible subsets S, and for variable i, calculate the difference in the model output before and after adding subset S:

[0031] φ i φ(S) = v(S ∪ {i}) - v(S) (11)

[0032] In the formula, v(S) is the evaluation function of subset S.

[0033] S52. Calculate the weighted average of the marginal contributions calculated for all subsets of each variable:

[0034]

[0035] In the formula, φ i is the SHAP value of variable i, representing its contribution to the model output; is the order weight of the variable appearance;

[0036] S53. Calculate the SHAP values for each variable on each sample, draw a variable contribution diagram, and then analyze the influencing factors of accident severity.

[0037] Furthermore, the Pearson correlation coefficient p is used to test the correlation between two continuous variables. The calculation method is as follows:

[0038]

[0039] In the formula, and are the values of the continuous variables and in the nth accident or the nth sample, respectively. and are the means of the continuous variables and , respectively.

[0040] Furthermore, the Cramer's V correlation coefficient is used to test the correlation between two discrete variables. The calculation method is as follows:

[0041]

[0042] In the formula, χ 2 is the chi-square statistic between the two discrete variables and . N is the number of accidents, K is the number of value types of the discrete variable , and R is the number of value types of the discrete variable .

[0043]

[0044] In the formula, N k is the number of accidents in which the discrete variable takes the value k, N r is the number of accidents in which the discrete variable takes the value r, and N kr is the frequency of the sample pair appearing.

[0045] Furthermore, the Spearman rank correlation coefficient ρ is used to test the correlation between a continuous variable and a discrete variable:

[0046]

[0047] In the formula, R i and are the ranks and average ranks of the continuous variable X C , respectively. S i and are the ranks and average ranks of the discrete variable X D , respectively.

[0048] Furthermore, in the SMOTENC algorithm process, for continuous variables, the value on the new sample is:

[0049]

[0050] In the formula, is the value of the j-th continuous variable on the i-th original sample, is the value of the j-th continuous variable of the randomly selected nearest neighbor sample, and δ j is a random number in the range [0, 1];

[0051] For discrete variables, the value on the new sample is:

[0052]

[0053] In the formula, N is the preset number of nearest neighbors, and Mode(·) represents taking the mode of the l-th discrete variable among n nearest neighbors

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0055] The present invention proposes a local cascade integration method that takes into account both the fitting degree of training data and the generalization ability of unknown data. By adopting the SMOTENC resampling technique, on the premise of considering the differences between discrete and continuous variables, the number of various accidents in the training set is balanced, and the model training effect and classification performance are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Att Figure 1 is the flow of the present invention;

[0057] Att Figure 2 is the Pearson correlation coefficient between continuous variables;

[0058] Att Figure 3 is the Cramer correlation coefficient between discrete variables;

[0059] Att Figure 4 is the Spearman rank correlation coefficient between continuous and discrete variables;

[0060] Att Figure 5 is the confusion matrix of the output result of the local cascade integration model;

[0061] Att Figure 6 is the contribution of each variable to fatal accidents;

[0062] Att Figure 7 is the contribution of each variable to injured accidents;

[0063] Att Figure 8 is the contribution of each variable to non-injury accidents. DETAILED DESCRIPTION

[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0065] To facilitate understanding of this embodiment, firstly, a method for analyzing factors affecting the severity of traffic accidents based on local cascade integration disclosed in an embodiment of the present invention is described. Figure 1 A flow chart of a method for analyzing factors affecting the severity of traffic accidents based on local cascade integration disclosed in an embodiment of the present invention is shown. Figure 1 As shown, the traffic accident severity influencing factor analysis method based on local cascade integration includes the following steps:

[0066] S1, obtain the accident data set D1, including driver, vehicle, road and environmental variable information;

[0067] S2, process the accident data set D1, remove some redundant attributes, divide it into training set and test set in proportion, balance the number of accidents of different severity (for example, divided into death, injury and no casualties) in the training set by SMOTENC algorithm, and obtain the accident data set D2; SMOTENC is a synthetic minority class oversampling technology for discrete and continuous variables;

[0068] S3, import the accident data set D2 into the local cascade integration model, adjust the hyperparameters of the local model through the Hyperopt method, and use k-fold cross validation to select the best hyperparameter combination. The Hyperopt method is a hyperparameter optimization method;

[0069] S4. Draw a confusion matrix based on the training results and select indicators to evaluate model performance;

[0070] S5. Use the machine learning output interpretation tool SHAP to visualize the model and analyze the factors affecting the severity of the accident.

[0071] Specifically, the specific method of S2 is:

[0072] S21. Remove accident data that has incomplete information or contains obvious errors;

[0073] S22. Select different correlation coefficients according to the variable type to test their correlation. If the correlation coefficient between two variables is ≥ 0.7, only one of the variables needs to be included in the modeling;

[0074] Among them, the Pearson correlation coefficient p is used to test the correlation between two continuous variables, and the calculation method is as follows:

[0075]

[0076] In the formula, and are the values of the continuous variables and in the nth accident or the nth sample respectively, and are the means of the continuous variables and respectively.

[0077] The Cramer's correlation coefficient V is used to test the correlation between two discrete variables, and the calculation method is:

[0078]

[0079] In the formula, χ 2 is the chi-square statistic between the two discrete variables and , N is the number of accidents, K is the number of value types of the discrete variable , and R is the number of value types of the discrete variable ;

[0080]

[0081] In the formula, N k is the number of accidents where the discrete variable takes the value k, N r is the number of accidents where the discrete variable takes the value r, and N kr is the frequency of occurrence of the sample pair .

[0082] The Spearman rank correlation coefficient ρ is used to test the correlation between a continuous variable and a discrete variable, and the calculation method is:

[0083]

[0084] In the formula, R i and are the rank and average rank of the continuous variable X C respectively, and S i and are the rank and average rank of the discrete variable X D respectively. Each observation in the variable will be replaced by its rank in the entire dataset. When there are identical values in the data, the ranks of these identical values will be averaged to avoid calculation biases caused by duplicate values.

[0085] S23. Divide the dataset D1 into a training set and a test set according to a ratio of 7:3, and use the SMOTENC algorithm to balance the number of accidents with different severities in the training set to obtain the accident dataset D2. The specific steps are as follows: First, identify the discrete variables in the training set; then, except for the samples with the largest proportion of severity (usually "no casualties"), for each sample of the remaining types of accidents, randomly select a similar sample in the k-nearest neighbors. For continuous variables, perform linear interpolation, and for discrete variables, take the mode of the nearest neighbors; finally, for different types of accidents, synthesize artificial data in an equal-proportion manner.

[0086] During the SMOTENC algorithm process, for continuous variables, the value on the new sample is:

[0087]

[0088] In the formula, is the value of the j-th continuous variable on the i-th original sample, is the value of the j-th continuous variable of the randomly selected nearest neighbor sample, and δ j is a random number in the range [0, 1];

[0089] For discrete variables, the value on the new sample is:

[0090]

[0091] In the formula, N is the preset number of nearest neighbors, and Mode(·) represents taking the mode of the l-th discrete variable among n nearest neighbors.

[0092] Among them, SMOTENC is an extended version of SMOTE (Synthetic Minority Over-sampling Technique), which is specifically used to process mixed datasets containing categorical features and continuous features.

[0093] Specifically, the specific method of S3 is as follows:

[0094] S31. The hyperparameters of the model include the maximum tree depth, the number of trees in the ensemble, the number of base learners, the maximum tree depth of the base learner, the learning rate of the base learner, and the minimum loss reduction when splitting nodes;

[0095] Hyperopt is an automatic hyperparameter tuning tool based on Bayesian optimization. The core idea is to construct a probability model of the objective function and select the hyperparameter combination through the Expected Improvement (EI) criterion;

[0096] S32. Randomly and evenly divide the accident dataset D2 into k = 5 subsets, where 4 subsets are used for training and the remaining 1 subset is used as a validation set to test the model performance. Repeat this process 5 times until the training is completed to find the optimal combination to maximize the model performance.

[0097] Further, the evaluation metrics selected in S4 include Accuracy, Recall, Precision, and F1-score, and their calculation formulas are as follows:

[0098]

[0099]

[0100]

[0101]

[0102] In the formula, TP is the true positive class, FN is the false negative class, FP is the false positive class, and TN is the true negative class.

[0103] Further, in S5, the KernelExplainer tool in the SHAP method is used to explain the output results of the local cascade ensemble model. The purpose of calculating the SHAP value is to fairly distribute the contribution (or influence) of each variable in the model to the output results. Its core idea is to regard the output value as the result of the collaborative work of all variables, and to distribute the importance by calculating the weighted average of the marginal contributions of each variable under different subset combinations, and sort according to the average value of the absolute SHAP value, and draw a variable contribution graph. The specific method of S5 is as follows:

[0104] S51. Select the average value of the training data as the baseline value, that is, the output of the model when there is no variable information; for the input variable set N = {1, 2,..., M}, calculate all possible subsets S, and for variable i, calculate the difference in the model output before and after the addition of subset S:

[0105] φ i (S) = v(S ∪ {i}) - v(S) (11)

[0106] In the formula, v(S) is the evaluation function of subset S;

[0107] S52. Calculate the weighted average of the marginal contributions calculated for all subsets of each variable:

[0108]

[0109] In the formula, φ i is the SHAP value of variable i, representing its contribution to the model output; is the order weight of the variable appearance;

[0110] S53. Calculate the SHAP value of each variable on each sample, draw a variable contribution graph, and then analyze the factors affecting the severity of the accident.

[0111] In order to more clearly demonstrate this method, a specific description is given below in conjunction with Example 1:

[0112] A method for analyzing factors affecting the severity of traffic accidents based on local cascade integration, such as Figure 1 As shown, the following steps are included:

[0113] S1 and accident dataset D1 are derived from 9,229 traffic accident records from 2011 to 2019 on multiple highways including Suiman Expressway, Harbin-Tongcheng Expressway, Tongsan Expressway, and Heha Expressway. The dataset records in detail the severity, time and location, cause and form of each accident, as well as information such as drivers and passengers, vehicles, roads, and environment.

[0114] S2, process the accident data set D1, remove some redundant attributes, divide it into training set and test set in proportion, balance the number of accidents of different severity (no casualties, injuries and deaths) in the training set through SMOTENC algorithm, and obtain the accident data set D2. The specific process is as follows:

[0115] S21. After removing the accident data with incomplete information or obvious errors, a total of 7946 accident samples were obtained that could be used for modeling research;

[0116] S22. Give variables clear and understandable names to improve code readability and subsequent analysis efficiency. Discrete variables include accident time (T), visibility (V), road alignment (GA), road surface condition (RS), season (S), accident type (CT), number of vehicles involved (NV), and whether seat belts were fastened (SB); continuous variables include the minimum age of the driver (AD_Y), the maximum age of the driver (AD_O), the minimum driving experience of the driver (DE_Y), and the maximum driving experience of the driver (DE_O).

[0117] Different correlation coefficients (value range [0,1]) were selected according to the variable type to test their correlation. If the correlation between two variables was strong (correlation coefficient ≥ 0.7), only one of the variables needed to be included in the modeling. In addition, if a discrete variable was divided into n attributes, in order to avoid complete collinearity between variables, n-1 dummy variables and 1 reference variable were required, and only n-1 dummy variables were included in the model. The statistical distribution characteristics of discrete variables and continuous variables are shown in Tables 1 and 2, respectively.

[0118] Table 1 Statistical distribution characteristics of discrete variables

[0119]

[0120]

[0121] Note: The variables in the brackets are abbreviations of variables during modeling; those marked with "#" are reference variables.

[0122] Table 2 Statistical Distribution Characteristics of Continuous Variables

[0123]

[0124] Note: For single-vehicle accidents, the minimum and maximum ages of the drivers are equal, and the minimum and maximum driving years are also equal.

[0125] S23. Divide the accident dataset D1 according to the ratio of 7:3 to obtain a training set and a test set with the number of samples being 5562 and 2384 respectively. Balance the number of accidents with different severities (no casualties, injured, and dead) in the training set through the SMOTENC algorithm, synthesize artificial data at a ratio of 1:1:1, and obtain an accident dataset D2 containing 12,909 samples.

[0126] S3. Import the accident dataset D2 into the local cascade integration model, adjust the hyperparameters of the local cascade integration model through the Hyperopt method, and use k-fold cross-validation to select the best hyperparameter combination. The specific process is as follows:

[0127] S31. The model hyperparameters include the maximum tree depth, the number of trees in the ensemble, the number of base learners, the maximum tree depth of the base learners, the learning rate of the base learners, and the minimum loss reduction when splitting nodes. Hyperopt is a hyperparameter automatic tuning tool based on Bayesian optimization. The core idea is to construct a probability model of the objective function and select the hyperparameter combination through the Expected Improvement (EI) criterion. First, evaluate the initial performance by randomly sampling a few hyperparameter combinations; then, divide the historical results into a high-performance group and a low-performance group through a tree-structured Parzen estimator, estimate the probability distributions of the hyperparameters respectively, with the goal of finding the parameters that make the probability density of the high-performance group high and the probability density of the low-performance group low, and select the next candidate parameter through the expected improvement criterion; finally, add the performance results of the new parameters to the historical data and iterate the optimization until the maximum number of trials is reached.

[0128] S32. Randomly and evenly divide the accident dataset D2 into k = 5 subsets, where 4 subsets are used for training and the remaining 1 subset is used as a validation set to test the model performance. Repeat this process 5 times until the training ends to find the optimal combination, as shown in Table 3.

[0129] Table 3 Optimal Results of Hyperparameters

[0130]

[0131] S4. Draw a confusion matrix based on the training results of the local cascade integration model, as Figure 5 shown. Here, 1, 2, and 3 represent non-injury accidents, injury accidents, and fatal accidents respectively, indicating that the model has high accuracy. The selected evaluation indicators include accuracy, recall rate, precision rate, and F1 score. According to their calculation formulas, the performance of the local cascade integration model is evaluated, and the results are shown in Table 4.

[0132] Table 4 Performance Evaluation Results of the Local Cascade Integration Model

[0133]

[0134] S5. Use the KernelExplainer tool in the SHAP method to explain the output results of the local cascade integration model, sort them according to the mean absolute value of SHAP, and draw a variable contribution diagram. The purpose of calculating SHAP values is to fairly distribute the contribution of each variable in the model to the output results. Its core idea is to regard the output value as the result of the collaborative work of all variables. By calculating the marginal contribution of each variable under different subset combinations and distributing the importance through weighted averaging. The specific process is as follows: Select the sample mean of the accident dataset D2 as the reference value, initialize the KernelExplainer. For each variable, calculate the SHAP value of each sample, sort the mean (i.e., importance) of the absolute value of SHAP of each variable, and obtain the variable contribution diagram. The horizontal axis represents the SHAP value, and a positive (negative) SHAP value indicates that the influencing factor has a positive (negative) impact on the dependent variable.

[0135] Figures 6 to 8 They are the contribution diagrams of each variable to fatal accidents, injury accidents, and non-injury accidents respectively. Based on this, the impact of each variable on the severity of accidents can be analyzed. For example, the variable "Good visibility (V_G)" has an overall positive SHAP value in Figure 6 while having a negative SHAP value in the other two diagrams, indicating that compared with the reference variable "Very poor visibility (V_EP)" (see Table 1), the improvement of visibility will increase the probability of fatal accidents and at the same time reduce the probability of injury and non-injury accidents (Note: The sum of the probabilities of the three accident severities is always 1). The potential reason is that when visibility is high, the vehicle speed is generally faster. Once an accident occurs, its severity is generally higher than that of an accident occurring at a low speed.

[0136] The above is only a preferred embodiment of the present invention, and it is not any other form of limitation to the present invention. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope protected by the present invention.

Claims

1. A method for analyzing factors affecting the severity of traffic accidents based on local cascade integration, characterized in that The following steps are involved: S1, obtain the accident data set D1, including driver, vehicle, road and environmental variable information; S2, process the accident data set D1, remove some redundant attributes, divide it into a training set and a test set in proportion, and balance the number of accidents of different severity in the training set through the SMOTENC algorithm to obtain the accident data set D2; S3, import the accident data set D2 into the local cascade integration model, adjust the hyperparameters of the local model through the Hyperopt method, and use k-fold cross validation to select the best hyperparameter combination; S4. Draw a confusion matrix based on the training results and select indicators to evaluate model performance; S5. Use the machine learning output interpretation tool SHAP to visualize the model and analyze the factors affecting the severity of the accident.

2. The method for analyzing influencing factors of traffic accident severity based on local cascade integration according to claim 1, wherein The specific method of S2 is: S21. Remove accident data that has incomplete information or contains obvious errors; S22. Select different correlation coefficients according to the variable type to test their correlation. If the correlation coefficient between two variables is ≥ 0.7, only one of the variables needs to be included in the modeling; Among them, Pearson's correlation coefficient p was used to test the correlation between two continuous variables; Cramer's correlation coefficient V was used to test the correlation between two discrete variables; The Spearman rank correlation coefficient ρ was used to test the correlation between continuous and discrete variables; S23. Divide the data set D1 into a training set and a test set in a ratio of 7:3, and use the SMOTENC algorithm to balance the number of accidents of different severity in the training set to obtain the accident data set D2. The specific steps are: first, clarify which discrete variables are in the training set; then, except for the sample with the largest severity (generally "no casualties"), for each of the remaining samples, randomly select a similar sample from the k nearest neighbors, perform linear interpolation for continuous variables, and take the mode of the nearest neighbors for discrete variables; finally, for different types of accidents, synthesize artificial data in a certain proportion.

3. The method for analyzing factors influencing the severity of traffic accidents based on local cascade integration according to claim 1, wherein The specific method of S3 is: S31, the hyperparameters of the model include maximum tree depth, number of trees in the ensemble, number of base learners, maximum tree depth of base learners, learning rate of base learners, and minimum loss reduction when splitting nodes; S32. Randomly and evenly divide the accident data set D2 into k=5 subsets, 4 of which are used for training and the remaining 1 is used as a validation set to test the model performance. Repeat this process 5 times until the training is completed and find the optimal combination to maximize the model performance.

4. The method for analyzing influencing factors of traffic accident severity based on local cascade integration according to claim 1, wherein The evaluation indicators selected in S4 include accuracy, recall, precision and F1 score, and the calculation formula is: Where TP is the true positive class, FN is the false negative class, FP is the false positive class, and TN is the true negative class.

5. The method for analyzing influencing factors of traffic accident severity based on local cascade integration according to claim 1, wherein The specific method of S5 is: S51. Select the average value of the training data as the benchmark value, that is, the output of the model without any variable information; for the input variable set N = {1, 2, ..., M}, calculate all possible subsets S, and for variable i, calculate the difference in model output before and after the subset S is added: φ i (S) = v(S ∪ {i}) - v(S) (11) Where v(S) is the evaluation function of subset S; S52. Take a weighted average of the marginal contributions calculated for all subsets of each variable: where φ i is the SHAP value of variable i, representing its contribution to the model output; is the sequential weight of the variable; S53. Calculate the SHAP values for each variable on each sample, draw a variable contribution diagram, and then analyze the influencing factors of accident severity.

6. The method for analyzing factors influencing the severity of traffic accidents based on local cascade integration according to claim 2, wherein: Use the Pearson correlation coefficient p to test the correlation between two continuous variables. The calculation method is as follows: In the formula, and are the values of the continuous variables and at the nth accident or the nth sample, and are the means of the continuous variables and respectively.

7. The method for analyzing the influencing factors of traffic accident severity based on local cascade integration according to claim 6, characterized in that: Use the Cramer's V correlation coefficient to test the correlation between two discrete variables. The calculation method is as follows: where χ 2 is the chi-square statistic between two discrete variables and , N is the number of accidents, K is the number of value types of the discrete variable X1 D , and R is the number of value types of the discrete variable ; where N k is a discrete variable and represents the number of accidents with a value of k, N r is a discrete variable and represents the number of accidents with a value of r, N kr is the frequency of occurrence of the sample pair appearing.

8. The method for analyzing factors influencing the severity of traffic accidents based on local cascade integration according to claim 7, wherein: Use the Spearman rank correlation coefficient ρ to test the correlation between a continuous variable and a discrete variable: where R i and are the rank and average rank of the continuous variable X C respectively, and S i and are the rank and average rank of the discrete variable X D respectively.

9. The method for analyzing influencing factors of traffic accident severity based on local cascade integration according to claim 2, wherein In the SMOTENC algorithm process, for continuous variables, the value on the new sample is: wherein, is the value of the j-th continuous variable on the i-th original sample, is the value of the j-th continuous variable of the randomly selected nearest neighbor sample, and δ j is a random number in the value range [0, 1]; For discrete variables, the value on the new sample is: In the formula, N is the preset number of nearest neighbors, and Mode(·) represents taking the mode of the l-th discrete variable among n nearest neighbors.

Citation Information

Cited By

  • TBM construction core database construction method and system

    CN120763148A

  • Maritime accident prediction method and device based on interpretable integrated machine learning

    CN121457753A

  • Maritime accident prediction method and device based on interpretable ensemble machine learning

    CN121457753B