Traffic accident severity prediction method and system based on multi-stage fusion

Through multi-stage fusion data processing and model prediction methods, the data imbalance and multi-source heterogeneous data problems in traffic accident severity prediction are solved, and high-precision traffic accident severity prediction is achieved, which improves the effects of traffic safety management and driver safety warning.

CN120296560AActive Publication Date: 2025-07-11CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510417958.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-11
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing traffic accident severity prediction methods have difficulties in dealing with multi-source heterogeneous data and data imbalance, resulting in insufficient prediction accuracy and reliability. A single model is susceptible to noise data and is difficult to meet actual needs.

Method used

A multi-stage fusion method is adopted, including data acquisition and preprocessing, preliminary prediction and feature fusion of logistic regression models, and quadratic prediction of AdaBoost model. Through undersampling and oversampling equilibrium data distribution, numerical feature standardization and categorical variable encoding, combined with Pearson correlation analysis and polynomial transformation, a strong classifier is built to improve prediction accuracy.

Benefits of technology

It realizes high-precision prediction of the severity of traffic accidents, improves the stability and generalization capabilities of the model, reduces the impact of noise data, improves the accuracy and reliability of predictions, and supports traffic management and driver safety warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296560A_ABST
    Figure CN120296560A_ABST
Patent Text Reader

Abstract

The invention provides a traffic accident severity prediction method and system based on multi-stage fusion, and relates to the field of traffic accident analysis and intelligent traffic. Comprising the following steps: collecting multi-source data related to a traffic accident, carrying out code conversion on classified variables after cleaning and processing data imbalance, and carrying out standardization processing on numerical characteristics; constructing a logistic regression model to perform preliminary prediction and output the probability of a corresponding category, performing correlation analysis on the accident severity by using historical data, selecting features with relatively high correlation to perform polynomial transformation, and fusing the transformed features and the result probability of preliminary prediction into new features; training the fused features by using the AdaBoost model again, and then predicting the features of the test set to obtain a final fusion prediction result; and finally, evaluating and optimizing the whole model. According to the method, various data resources are integrated in a multi-stage fusion mode, the fused model can integrate the advantages of a single model, and higher prediction accuracy is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of traffic accident analysis and intelligent transportation, and particularly relates to a method and system for predicting traffic accident severity based on multi-stage fusion. Background Art

[0002] The occurrence and severity of traffic accidents are comprehensively affected by a variety of complex factors. Among geographical location factors, the traffic conditions of different terrains and regions such as urban roads, rural roads, and mountain roads vary significantly, and the probability and severity of accidents also differ; in terms of road conditions, the flatness, width, curve curvature of the road, and the perfection degree of traffic signs will all have an important impact on vehicle driving safety; among environmental factors, weather conditions such as heavy rain, heavy snow, fog, strong wind and other bad weather will greatly reduce the visibility of the road and increase the difficulty of vehicle control; vehicle driving states, including speed, acceleration, steering, etc., and driver behaviors such as fatigue driving, illegal driving, and distracted attention are also important factors leading to accidents and affecting accident severity; in addition, the magnitude and variation law of traffic flow will also indirectly affect the occurrence frequency and severity of accidents.

[0003] In the field of traffic accident research, accurately predicting accident severity is one of the key tasks. Traditional prediction methods mainly rely on simple statistical analysis or a single model architecture. However, the occurrence mechanism of traffic accidents is extremely complex and is affected by the interaction of multiple factors. Facing such a complex situation of multi-factor interweaving, traditional methods are difficult to comprehensively capture the potential laws and features in the data, resulting in the prediction accuracy being difficult to meet the actual needs.

[0004] With the rapid development of information technology, the advent of the big data era has brought new opportunities for traffic accident prediction. Machine learning technology, with its powerful data analysis and pattern recognition capabilities, has achieved remarkable results in many fields and also provides new ideas for predicting traffic accident severity. However, existing machine learning models still face many challenges when applied to traffic accident prediction.

[0005] First of all, traffic accident data has high complexity and diversity, covering various aspects of information such as accident occurrence time, location, vehicle model, speed, driving trajectory, weather conditions, road conditions, etc. These data are significantly different in format, scale, and nature, forming multi-source heterogeneous data. Some existing models are difficult to effectively integrate and mine the internal relationships between different types of data when processing this multi-source heterogeneous data, resulting in a waste of a large amount of valuable information and seriously affecting prediction accuracy.

[0006] Secondly, the problem of data imbalance is also an important factor that restricts the performance of existing models. In actual traffic accident data sets, the distribution of accident severity is usually extremely uneven. The proportion of samples of minor accidents is often very high, while samples of moderate and severe accidents are relatively scarce. This serious data imbalance makes it easy for the model to over-focus on majority class samples during training, while the learning and recognition capabilities of minority class samples are insufficient. When the model is applied to actual predictions, the prediction effect for minority class samples is poor, and it is impossible to accurately identify and judge accidents with higher degrees of harm, which greatly reduces the reliability and effectiveness of the model in predicting the overall accident severity.

[0007] In addition, the balance between model complexity and generalization ability is also a difficult problem that needs to be solved urgently. In order to pursue higher fitting accuracy, some models have built overly complex structures. Although they perform well on training data, they are prone to overfitting when faced with new data, resulting in a sharp drop in performance in test sets or actual applications. On the contrary, although some simple models have good generalization ability, they have limited ability to learn and express data, and cannot fully capture the complex characteristics and laws in traffic accident data, resulting in prediction accuracy that is difficult to meet actual needs.

[0008] Chinese patent document CN108710967A discloses a highway traffic accident severity prediction method based on data fusion and support vector machine. Specifically, multiple types of data are collected first, and then the variable factors of the data samples are reduced in dimension and normalized. The traffic accident severity prediction model is constructed using the support vector machine algorithm, and the variable factor vector with the predicted accident after dimension reduction is brought into the prediction model for prediction. This scheme predicts the severity of traffic accidents based on data fusion and uses the support vector machine method. Under normal circumstances, there is a serious data imbalance in the severity of accidents, and the sources of influencing factors are wide-ranging. However, this scheme does not perform related operations such as data imbalance processing and feature selection, and only performs one prediction, which is easily interfered by noise data and has a single degree of overfitting.

[0009] In summary, given the high incidence of traffic accidents and the limitations of existing prediction methods, there is an urgent need for an innovative and efficient traffic accident severity prediction method. Summary of the invention

[0010] The technical problem to be solved by the present invention is to provide a traffic accident severity prediction method and system based on multi-stage fusion in view of the shortcomings of the existing technology, which can achieve high-precision prediction of the severity of traffic accidents.

[0011] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0012] In a first aspect, the present invention provides a traffic accident severity prediction method based on multi-stage fusion, comprising the following steps:

[0013] S1. Data collection and preprocessing: Collect multi-source data related to traffic accidents, and clean the data to remove duplicate and invalid data; Use undersampling and oversampling to balance the data distribution of accident severity; Encode and transform categorical variables and standardize numerical features to make the data meet the requirements of model input;

[0014] S2. Preliminary prediction and feature fusion: Construct a logistic regression model for preliminary prediction and output probability results. At the same time, use historical data to conduct a correlation analysis on accident severity, select features with higher correlation for polynomial transformation, and fuse the transformed features and the probability results of preliminary prediction into new features;

[0015] S3. Secondary prediction: Use the AdaBoost model to train the fused features again, and then predict the test set to obtain the final fused prediction result;

[0016] S4. Result evaluation and model optimization: Use accuracy, recall, and F1 value metrics to draw ROC and PR curves for the fused model to evaluate the prediction results.

[0017] Furthermore, the multi-source data related to traffic accidents includes accident record data, weather condition data, and road monitoring data.

[0018] Furthermore, the accident record data includes data information such as the exact time of the accident, the location accurately positioned by longitude and latitude coordinates, the exact model and license plate number of the involved vehicle, and the specific situation of casualties. Generally, detailed accident records are obtained through cooperation with the traffic management department.

[0019] Furthermore, the weather condition data includes data such as temperature, humidity, wind speed, precipitation, and visibility at the time of the accident. Generally, it is obtained through the professional data interface of the meteorological department.

[0020] Furthermore, the road monitoring data includes data such as the road surface condition and traffic flow of the road where the accident occurred. Generally, it is collected by monitoring equipment beside the road.

[0021] Furthermore, the specific meaning of using undersampling and oversampling to balance the data distribution of accident severity in step S1 is as follows: Conduct a comprehensive statistical analysis on the multi-source data related to traffic accidents after cleaning, and intuitively present the sample quantity distribution of different severity categories through a histogram or pie chart; If it is found that the data is seriously unbalanced, the following processing is carried out:

[0022] For the minor accident samples in the majority category, use the undersampling method to effectively reduce their absolute quantity in the dataset and reduce their dominant influence on model training;

[0023] For a small number of severe accident samples, an oversampling algorithm is used to generate new synthetic samples in the feature space of the minority-class samples to increase their sample size.

[0024] This processing can make the samples of different severities reach a relatively balanced distribution state in the dataset, ensuring that the model can fully learn the characteristic patterns of various accidents during the training process, and avoiding insufficient recognition and prediction capabilities of the model for minority-class samples due to data imbalance.

[0025] Furthermore, the categorical variable encoding conversion in step S1 refers to converting categorical variables into consecutive integer encodings starting from 0. For example, encoding "sunny day" as 0, "rainy day" as 1, "snowy day" as 2, etc. In this way, the model can correctly recognize and utilize this categorical information when processing data, avoiding affecting the prediction results due to improper representation of categorical variables.

[0026] Furthermore, the standardization processing in step S1 is obtained through calculating the mean and standard deviation of each numerical feature and using the conversion formula:

[0027]

[0028] where: X new is the standardized feature value, x is the original feature value, μ is the mean of this feature, and σ is the standard deviation.

[0029] After standardization processing, all numerical features will be converted into a standard normal distribution with a mean of 0 and a standard deviation of 1, making different features comparable in numerical scale, avoiding the model over-focusing on or ignoring certain features during the training process due to excessive differences in feature dimensions, improving the efficiency and stability of model training, and helping the model converge to the optimal solution faster.

[0030] Furthermore, in step S2, the logistic regression model outputs the prediction probability through the One vs Rest strategy, that is, creating K binary classification models, where K represents the number of accident severity categories, usually set to three categories: minor, moderate, and severe. Each model takes one category as the positive class and the remaining categories as the negative class, which is represented by the sigmoid function, that is:

[0031]

[0032] where: σ(z) represents the sigmoid function, and the output range is between 0 and 1, representing the probability that the sample belongs to the positive class, and z is the linear combination, that is:

[0033] z = θ0 + θ1x1 + θ2x2 + …… + θ n x n = θ T x

[0034] θ is the parameter vector of the model, T represents transpose, and n represents the number of features x.

[0035] Furthermore, the training objective of the logistic regression model in step S2 is to find the optimal parameter θ to maximize the likelihood function. In the solution of the present invention, the logarithmic likelihood function is used and solved by optimization algorithms such as gradient descent. The formula for the logarithmic likelihood function L(θ) is:

[0036]

[0037] where m is the number of samples, y (i) is the true label of the i-th sample, x (i) is the feature vector of the i-th sample, and σ(θ T x (i) ) is the probability that the i-th sample belongs to the positive class after being calculated by the sigmoid function.

[0038] Furthermore, in step S2, through correlation analysis, the absolute value of the Pearson correlation coefficient between each feature and the target variable "accident severity" is calculated, and the features with relatively high correlation are screened out. The formula for the Pearson correlation coefficient (r) is:

[0039]

[0040] where x i and y i are respectively the i-th sample values of the feature x and the target variable y, and are respectively the means of the feature x and the target variable y; n represents the number of the feature x and the target variable y.

[0041] After that, a quadratic polynomial transformation is performed on the variables with relatively high correlation to generate multiple new features, and at the same time, the predicted probability output by the logistic regression model is incorporated as a new feature.

[0042] Furthermore, in step S3, AdaBoost iteratively trains multiple weak classifiers and combines them into a strong classifier. In each round of iteration, AdaBoost adjusts the weights of the samples according to the error rate of the previous round of classifier. The weight α i of the weak classifier h t (x) in the t-th round of iteration is calculated by the formula:

[0043]

[0044] where ∈ t is the error rate of the weak classifier in the t-th round of iteration, and the sample weight update formula is:

[0045]

[0046] Among them, represents x i the weight of the t-th round of iteration, represents x i the weight after the t-th round of iteration update; h t (x i ) represents the prediction result of the weak learner; exp represents the exponential function, and Z t represents the normalization factor, which is used to ensure that the sum of the weights of all samples is 1.

[0047] Furthermore, the final model strong classifier H(x) in step S3 is:

[0048]

[0049] Among them, N is the number of weak classifiers, that is, the total number of rounds of iteration of the AdaBoost algorithm, and sign is the sign function, which outputs 1 when the input value is greater than 0 and -1 when it is less than 0.

[0050] In a second aspect, the present invention also provides a traffic accident severity prediction system based on multi-stage fusion, including a processor and a memory, wherein computer program code instructions are stored on the memory;

[0051] When the computer program code instructions are called by the processor, the processor is caused to execute the traffic accident severity prediction method based on multi-stage fusion as described above.

[0052] Advantages of the present invention:

[0053] The traffic accident severity prediction method provided by the present invention first cleans the collected multi-source data related to traffic accidents to remove duplicate and invalid data; uses undersampling and oversampling to balance the accident severity data distribution; encodes and converts categorical variables and normalizes numerical features so that the data can meet the model input requirements. Among them, the handling of data imbalance can ensure that the model can fully learn the characteristic patterns of various accidents during the training process, avoid the insufficient recognition and prediction ability of the model for minority class samples due to data imbalance, and thus ensure the accuracy of model prediction; through the encoding and conversion of categorical variables, the model can correctly recognize and utilize these categorical information when processing data, avoiding the influence on the prediction result due to the improper representation of categorical variables; also through the normalization of numerical features, different features are comparable on the numerical scale, avoiding the model's over-concern or neglect of certain features during the training process due to excessive differences in feature dimensions, improving the efficiency and stability of model training, and helping the model converge to the optimal solution faster. Then, a logistic regression model is constructed for preliminary prediction and the probability of the corresponding category is output. At the same time, the historical data is used to conduct a correlation analysis on the accident severity, and the features with higher correlations are selected for polynomial transformation, and the transformed features and the result probability of the preliminary prediction are fused into new features. After that, the AdaBoost model is used to train the fused features, and then the features of the test set are predicted to obtain the final fused prediction result. Finally, the entire model is evaluated and optimized.

[0054] In the prediction of traffic accident severity, due to the large variety and quantity of data, if data imbalance processing, feature selection and other operations are not carried out and single model prediction is directly performed, overfitting of a single severity is likely to occur. At the same time, a single prediction model is greatly affected by noise data and is difficult to identify the associations between multiple features. Therefore, a single logistic regression model is easily affected by noise data. The solution of the present invention first processes the collected multi-source data related to traffic accidents, such as data imbalance and data cleaning, and then constructs a logistic regression model for preliminary prediction. At the same time, the features with higher correlations are selected for polynomial transformation, and the obtained features and the result probability of the preliminary prediction are fused into new features and input into the secondary prediction model. The secondary prediction model uses the AdaBoost model to train the fused features, and then predicts the features of the test set to obtain the final fused prediction result. Through the secondary iteration by AdaBoost, the noise influence brought by the large variety and quantity of data in the prediction of a single logistic regression model can be further reduced. The present invention effectively integrates various data resources through a multi-stage fusion method, and the fused model effectively integrates the advantages of single models, improving the prediction accuracy compared with single models.

[0055] The present invention can overcome the deficiencies of the prior art, achieve high-precision prediction of traffic accident severity, provide strong technical support for traffic management departments to formulate scientific and reasonable traffic policies, optimize traffic resource allocation, and provide timely and accurate safety warnings for drivers, thereby improving the overall traffic safety level and reducing the losses caused by traffic accidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0057] Figure 1 It is a schematic flowchart of the basic process of the traffic accident severity prediction method based on multi-stage fusion provided in Embodiment 1 of the present invention.

[0058] Figure 2 It is a schematic structural diagram of the logical regression model in the embodiment of the present invention.

[0059] Figure 3 It is a schematic structural diagram of the AdaBoost model in the embodiment of the present invention.

[0060] Figure 4 It is an ROC curve graph for evaluating the prediction results in the embodiment of the present invention.

[0061] Figure 5 It is a PR curve graph for evaluating the prediction results in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The following further illustrates the invention in conjunction with the embodiments and the drawings, but does not limit the scope of the present invention.

[0063] Embodiment 1

[0064] As Figure 1 shown, this embodiment provides a traffic accident severity prediction method based on multi-stage fusion, including the following steps:

[0065] S1. Data collection and preprocessing: Collect multi-source data related to traffic accidents, remove duplicate and invalid data through cleaning; balance the data distribution of accident severity using undersampling and oversampling; encode and transform categorical variables and standardize numerical features to make the data meet the model input requirements. Specifically as follows:

[0066] By cooperating with the traffic management department to obtain its detailed accident records, which cover information such as the exact time of the accident, the location accurately positioned by longitude and latitude coordinates, the exact models and license plate numbers of the vehicles involved, and the specific situation of casualties; obtaining the weather conditions such as temperature, humidity, wind speed, precipitation, visibility, etc. at the time of the accident through the professional data interface with the meteorological department; using the monitoring equipment beside the road to collect information such as the road surface conditions and traffic flow of the road.

[0067] Table 1 Data Collection and Classification

[0068]

[0069] Since the data collection process is vulnerable to various factors, some data may be missing, incorrect, or invalid. For example, vehicle sensors may transmit incorrect data or interrupt data transmission due to bad weather, electromagnetic interference, or equipment failure, resulting in the absence of some key data points; manual records of the traffic management department may have clerical errors or incomplete information entry; there may also be differences in the data formats and standards of different data sources. For instance, vehicle sensor data may adopt a specific binary coding format, while meteorological department data may be in the form of text reports, all of which need to be uniformly processed. Therefore, in some embodiments of the present invention, a comprehensive data cleaning operation is first carried out on the collected data. The data is strictly screened according to the internal logical relationship and reasonable range of the data. For duplicate data records, deduplication can be performed according to the unique identifier of the data, such as the combination of accident number, vehicle frame number, and timestamp.

[0070] After completing the data cleaning, it is necessary to further address the problem of data imbalance. First, a comprehensive statistical analysis is carried out on the accident severity data, and a bar chart or pie chart is drawn to visually present the distribution of the sample numbers of different severity categories. If it is found that the data is severely imbalanced, for example, in an embodiment of the present invention, the proportion of minor accident samples is too high, such as exceeding 90% in some embodiments, while the proportion of severe accident samples is relatively scarce, such as less than 5% in some embodiments, corresponding measures need to be taken. For the majority category of minor accident samples, the undersampling method can be used to effectively reduce their absolute numbers in the dataset and reduce their dominant influence on model training; for the minority category of severe accident samples, oversampling algorithms can be used to generate new synthetic samples in the feature space of the minority samples, increasing their sample numbers and enabling the samples of different severities to reach a relatively balanced distribution state in the dataset, ensuring that the model can fully learn the feature patterns of various accidents during the training process and avoiding insufficient recognition and prediction capabilities of the model for minority samples due to data imbalance.

[0071] Subsequently, encoding transformation is performed on categorical variables. Such as road types (expressways, urban arterial roads, rural paths, etc.), weather conditions (sunny, rainy, snowy, foggy, etc.), etc., and these categorical variables are converted into consecutive integer encodings starting from 0. For example, "sunny" is encoded as 0, "rainy" is encoded as 1, "snowy" is encoded as 2, etc. In this way, the model can correctly identify and utilize this categorical information when processing data, avoiding affecting the prediction results due to improper representation of categorical variables.

[0072] Finally, standardization processing is performed on numerical features. Such as the distance information after numericalization of the longitude and latitude coordinates of the accident location, vehicle driving speed, acceleration value, temperature and humidity values at the time of the accident, etc. By calculating the mean and standard deviation of each numerical feature, through the transformation formula:

[0073]

[0074] where X new is the standardized feature value, x is the original feature value, μ is the mean of this feature, and σ is the standard deviation. After standardization processing, all numerical features will be converted into a standard normal distribution with a mean of 0 and a standard deviation of 1, making different features comparable on the numerical scale, avoiding the model over-focusing on or ignoring certain features during the training process due to excessive differences in feature dimensions, improving the efficiency and stability of model training, and helping the model converge to the optimal solution faster.

[0075] S2. Preliminary prediction and feature fusion: Construct a logistic regression model for preliminary prediction and output probability results. At the same time, use historical data to conduct a correlation analysis on the accident severity, select features with higher correlations for polynomial transformation, and fuse the transformed features and the probability results of the preliminary prediction into new features, specifically as follows:

[0076] S201. Model construction

[0077] In this stage, a logistic regression model is selected for preliminary prediction. Logistic regression is a commonly used linear classification model. When dealing with multi-classification problems, the One-vs-Rest (OvR) strategy is adopted here, and the model structure is as Figure 2 shown. For the accident severity prediction problem, assuming there are K categories of accident severity, K binary classification models will be constructed. Each binary classification model takes one of the categories as the positive class and the remaining K - 1 categories as the negative class, and is represented by the sigmoid function, that is:

[0078]

[0079] where: σ(z) represents the sigmoid function, and the output range is between 0 and 1, indicating the probability that the sample belongs to the positive class, and z is the linear combination, that is:

[0080] z = θ0 + θ1x1 + θ2x2 + …… + θ n x n = θ T x

[0081] θ is the parameter vector of the model, T represents transpose, and n represents the number of features x.

[0082] S202. Data Preparation and Input

[0083] Input the pre - processed data (including traffic accident - related data after cleaning, balancing, encoding, and standardization) into the logistic regression model. The dimensionality and format of the features of the pre - processed data need to match the input requirements of the logistic regression model. For example, if the data has n features after processing, then when inputting into the logistic regression, the dimensionality of the input vector for each sample is n.

[0084] S203. Model Training

[0085] Use the training dataset to train the logistic regression model. The training objective of the logistic regression model is to find the optimal parameter θ to maximize the likelihood function. Usually, the log - likelihood function is used and solved through optimization algorithms such as gradient descent. The formula for the log - likelihood function L(θ) is:

[0086]

[0087] where m is the number of samples, y (i) is the true label of the i - th sample, x (i) is the feature vector of the i - th sample, and σ(θ T x (i) ) is the probability that the i - th sample belongs to the positive class after being calculated by the sigmoid function.

[0088] S204. Correlation Analysis and Feature Fusion

[0089] Make a preliminary prediction through the logistic regression model and output the probability results. At the same time, use historical data to conduct a correlation analysis on the accident severity, select features with higher correlation by means of the Pearson correlation coefficient, and perform polynomial transformation. The formula for the Pearson correlation coefficient r is:

[0090]

[0091] where x i and y i are the i - th sample values of feature x and target variable y respectively, and are the means of feature x and target variable y respectively; n represents the number of feature x and target variable y;

[0092] After that, a quadratic polynomial transformation is performed on the variables with relatively high correlation, generating multiple new features again, and incorporating the predicted probabilities output by the logistic regression model as new features.

[0093] S3. Select the AdaBoost model for secondary prediction: Use the AdaBoost model to train the fused features again, and then predict the test set to obtain the final fused prediction result, as follows:

[0094] S301. Model selection and construction

[0095] Select the AdaBoost model as the model for secondary prediction. AdaBoost iteratively trains multiple weak classifiers and combines them into a strong classifier. The model structure is as Figure 3 shown. In each round of iteration, AdaBoost adjusts the weights of the samples according to the error rate of the previous round of classifier. The weight α i of the weak classifier h t (x) in the t-th round of iteration is calculated by the formula:

[0096]

[0097] where ∈ t is the error rate of the weak classifier in the t-th round of iteration, and the sample weight update formula is:

[0098]

[0099] where represents the weight of x i in the t-th round of iteration, represents the weight of x i after updating in the t-th round of iteration; h t (x i ) represents the prediction result of the weak learner; exp represents the exponential function, and Z t represents the normalization factor, which is used to ensure that the sum of all sample weights is 1.

[0100] S302. Data preparation and input

[0101] Input the dataset with new features added after preliminary prediction as the input data into the AdaBoost model. At this time, the data contains the variables with high correlation with "accident severity" after polynomial transformation and the result probabilities predicted by the logistic regression model. The dimension and feature information of the data are more abundant, which helps the AdaBoost model learn more complex patterns and relationships.

[0102] S303. Model training

[0103] Train the AdaBoost model using the fused features. During the training process, AdaBoost iteratively trains multiple weak classifiers and combines them into a strong classifier. The final strong classifier H(x) of the model is as follows:

[0104]

[0105] where N is the number of weak classifiers, i.e., the total number of iterations of the AdaBoost algorithm, sign is the sign function, which outputs 1 when the input value is greater than 0 and -1 when it is less than 0.

[0106] S4. Result evaluation and model optimization: Compare the accuracy, recall rate, F1 value metrics, and ROC and PR curves of the individual logistic regression model, AdaBoost model, and the fused model to evaluate the prediction results, as follows:

[0107] Conduct a comprehensive evaluation of the prediction results and use metrics such as accuracy, recall rate, and F1 value to measure the performance of the model. The formula for calculating accuracy is:

[0108]

[0109] The formula for calculating the recall rate (for a certain category) is:

[0110]

[0111] The formula for calculating the F1 value is:

[0112]

[0113] where TP represents true positive, TN represents true negative, FP represents false positive, FN represents false negative; Precision represents precision. By plotting the ROC curve with the false positive rate as the abscissa and the true positive rate as the ordinate, visually display the classification performance of the model at different thresholds; by plotting the PR curve with the recall rate as the abscissa and the precision as the ordinate, evaluate the precision performance of the model at different recall levels; select the key features of the combined model and add noise points with a certain multiple (0.8 - 1.2 times) to these features to visually observe the fluctuation of the model's accuracy and evaluate the stability of the model.

[0114] Embodiment 2

[0115] Based on the same inventive concept, this embodiment is a system embodiment corresponding to the above method Embodiment 1 and can be implemented in cooperation with the manner of the above Embodiment 1.

[0116] This embodiment provides a traffic accident severity prediction system based on multi-stage fusion, including a processor and a memory, where computer program code instructions are stored on the memory; when the computer program code instructions are called by the processor, the processor is caused to execute the steps of the traffic accident severity prediction method based on multi-stage fusion described in Embodiment 1 above.

[0117] The relevant technical details mentioned in the above Embodiment 1 are still valid in this embodiment, and the repeated parts will not be elaborated.

[0118] To verify the feasibility and effectiveness of the solution of the present invention, it is further illustrated by the following application embodiments.

[0119] Embodiment 3: Application Embodiment

[0120] First, collect the relevant data of road traffic accidents in the United States from January 15, 2016 to January 1, 2022, including many factors such as the accident location, driving time, vehicles involved in the accident, temperature, wind chill index, humidity, air pressure, visibility, wind direction, wind speed, precipitation, etc. Preprocess the collected data according to the method of Embodiment 1 above, such as encoding, normalization, undersampling, oversampling, etc. First, pass the original data through a logistic regression model to preliminarily predict the accident severity, calculate the absolute value of the Pearson correlation coefficient between each feature and the target variable "accident severity" through historical data, select the features with higher correlation, then fuse the preliminarily predicted results as the fused features, and finally perform a secondary prediction through the AdaBoost model. Evaluate the results of the fused model according to the method of Embodiment 1 above, and compare the results of the fused model with those of individual models (logistic regression, AdaBoost, and KNN, SVM, Naive Bayes). The results are as Figure 2 shown.

[0121] Table 2 Comparison Table of Model Effects

[0122]

[0123] As can be seen from the data in Table 2 above, the accuracy (Accuracy), recall rate (Recall), F1 value, and overall accuracy of the fused model of the present invention are improved compared with those of individual models. The ROC curve and PR curve of the fused model are respectively as Figure 4 and Figure 5 shown. Figure 4 In the ROC curve of , AUC (Area Under Curve) refers to the area enclosed by the ROC curve and the coordinate axes. The closer AUC is to 1.0, the higher the authenticity of the method. Figure 5In the PR curve, AP refers to the area enclosed by the PR curve and the coordinate axes. Generally, the higher the AP value, the better. The results show that the performance of the fused model of the present invention has been greatly improved compared with that of the individual models.

[0124] Obviously, the above embodiments are only preferred examples of the present invention and do not limit the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.

Claims

1. A traffic accident severity prediction method based on multi-stage fusion, characterized in that It includes the following steps: S1. Data collection and preprocessing: Collect multi-source data related to traffic accidents, and clean to remove duplicate and invalid data; Use undersampling and oversampling to balance the data distribution of accident severity; Encode and transform categorical variables and standardize numerical features to make the data meet the model input requirements; S2. Preliminary prediction and feature fusion: Construct a logistic regression model for preliminary prediction and output probability results. At the same time, use historical data to conduct a correlation analysis on accident severity, select features with higher correlations for polynomial transformation, and fuse the transformed features and the probability results of the preliminary prediction into new features; S3. Secondary prediction: Use the AdaBoost model to train the fused features again, and then predict the test set to obtain the final fused prediction result; S4. Result evaluation and model optimization: Use accuracy, recall rate, and F1 value metrics to draw ROC and PR curves for the fused model to evaluate the prediction results.

2. The method for predicting the severity of traffic accidents based on multi-stage fusion according to claim 1, wherein The multi-source data related to traffic accidents includes accident record data, weather condition data, and road monitoring data; The accident record data includes the exact time of the accident, the location accurately positioned by longitude and latitude coordinates, the accurate model and license plate number of the vehicle involved, and the specific situation of casualties; The weather condition data includes the temperature, humidity, wind speed, precipitation, and visibility at the time of the accident; The road monitoring data includes the road surface condition and traffic flow of the road where the accident occurred.

3. The traffic accident severity prediction method based on multi-stage fusion according to claim 1, wherein The specific meaning of using undersampling and oversampling to balance the data distribution of accident severity in step S1 is: Conduct a comprehensive statistical analysis on the multi-source data related to traffic accidents after cleaning, and intuitively present the sample quantity distribution of different severity categories through a bar chart or pie chart; If it is found that the data is seriously unbalanced, the following processing is done: For the minor accident samples of the majority class, use the undersampling method to effectively reduce their absolute quantity in the dataset and reduce their dominant influence on model training; For the severe accident samples of the minority class, use the oversampling algorithm to generate new synthetic samples in the feature space of the minority class samples to increase their sample quantity; The categorical variable encoding transformation in step S1 refers to converting the categorical variable into a continuous integer encoding starting from 0.

4. The traffic accident severity prediction method based on multi-stage fusion according to claim 1, characterized in that The standardization processing in step S1 is obtained through the following conversion formula by calculating the mean and standard deviation of each numerical feature: Where: X new is the standardized eigenvalue, x is the original eigenvalue, μ is the mean of this feature, and σ is the standard deviation.

5. The method for predicting the severity of traffic accidents based on multi-stage fusion according to claim 4, wherein In step S2, the logistic regression model outputs the prediction probability through the One vs Rest strategy, that is, create K binary classification models, where K represents the number of accident severity categories. Each model takes one category as the positive class and the remaining categories as the negative class, and is represented by the sigmoid function, that is: Among them: σ(z) represents the sigmiod function, and the output range is between 0 and 1, indicating the probability that the sample belongs to the positive class. z is a linear combination, that is: z = θ0 + θ1x1 + θ2x2 + …… + θ n x n = θ T x θ is the parameter vector of the model, T represents the transpose, and n represents the number of features x.

6. The method for predicting the severity of traffic accidents based on multi-stage fusion according to claim 5, characterized in that The training objective of the logistic regression model in step S2 is to find the optimal parameter θ to maximize the log-likelihood function, and it is solved through the gradient descent optimization algorithm; The formula for the log-likelihood function L(θ) is: where m is the number of samples, y (i) is the true label of the i-th sample, x (i) is the feature vector of the i-th sample, σ(θ T x (i) ) is the probability that the i-th sample belongs to the positive class after being calculated by the sigmoid function.

7. The method for predicting the severity of traffic accidents based on multi-stage fusion according to claim 6, characterized in that In step S2, through correlation analysis, the absolute value of the Pearson correlation coefficient between each feature and the target variable "accident severity" is calculated, and features with relatively high correlation are selected; the formula for the Pearson correlation coefficient r is: where x i and y i are the i-th sample values of the feature x and the target variable y, respectively, and are the means of the feature x and the target variable y, respectively; n represents the number of the feature x and the target variable y; After that, a quadratic polynomial transformation is performed on the variables with relatively high correlation to generate multiple new features, and at the same time, the predicted probability output by the logistic regression model is incorporated as a new feature.

8. The traffic accident severity prediction method based on multi-stage fusion according to claim 7, characterized in that, In step S3, AdaBoost iteratively trains multiple weak classifiers and combines them into a strong classifier. In each round of iteration, AdaBoost adjusts the weights of the samples according to the error rate of the previous round of classifier; Among them, the weight α of the weak classifier h i (x) in the t-th round of iteration t The calculation formula is as follows: where, ∈ t is the error rate of the weak classifier in the t-th round of iteration, and the sample weight update formula is: Among them, represents x i the weight of the t-th round of iteration, represents x i the weight after the t-th round of iteration update; h t (x i ) represents the prediction result of the weak learner; exp represents the exponential function, Z t represents the normalization factor, which is used to ensure that the sum of all sample weights is 1.

9. The method for predicting traffic accident severity based on multi-stage fusion according to claim 8, wherein The final model strong classifier H(x) in step S3 is: where N is the number of weak classifiers, that is, the total number of iterations of the AdaBoost algorithm, sign is the sign function, which outputs 1 when the input value is greater than 0 and -1 when it is less than 0.

10. A traffic accident severity prediction system based on multi-stage fusion, characterized in that, It includes a processor and a memory, where computer program code instructions are stored on the memory; When the computer program code instructions are called by the processor, the processor is caused to execute the above-mentioned method for predicting traffic accident severity based on multi-stage fusion.

Citation Information

Patent Citations

  • Traffic accident severity prediction method and system and computer storage medium

    CN116128108A

  • Marine traffic accident severity prediction method based on feature engineering

    CN119250257A