A method for predicting accident severity in mixed autonomous and manual driving situations

By constructing a joint prediction model in a mixed environment of autonomous driving and manual driving, the problems of insufficient data and unobserved heterogeneity are solved, the accuracy of accident severity prediction and the goodness of fit of the model are improved, and effective analysis of accident characteristics, road characteristics and built environment characteristics are achieved.

CN119904988BActive Publication Date: 2025-09-02CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510043489.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-09-02
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing accident severity prediction model in the mixed environment of autonomous driving and manual driving has problems such as insufficient data volume, missing key variables, missing data or unverified data, resulting in poor generalization capabilities of the model, unable to fully explore influencing factors, and ignore the interrelationship between accident severity between different types of vehicles and the unobserved heterogeneity in the data set.

Method used

By obtaining traffic accident data in a mixed environment of autonomous driving and manual driving, extracting road traffic characteristics and built environment characteristics, building a joint prediction model, using machine learning to screen the characteristic variables that affect severity, and using a random parameter bivariate Probit model to predict the accident severity, solving the problem of insufficient data and unobserved heterogeneity.

Benefits of technology

The accuracy of the prediction of accident severity and the goodness of fit of the model are improved, and the impact of accident characteristics, road characteristics and built environment characteristics on the severity of the accident are analyzed, which improves the prediction accuracy and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904988B_ABST
    Figure CN119904988B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of accident cause analysis and discloses a method for predicting the severity of accidents in mixed autonomous and manual driving environments. The method comprises obtaining traffic accident data in a mixed autonomous and manual driving environment; extracting road traffic characteristics and built environment characteristics at the accident location based on the accident location in the traffic accident data; matching the traffic accident data with the road traffic characteristics and built environment characteristics at the accident location to establish a dataset of characteristic variables for autonomous and manual driving traffic accidents; screening characteristic variables influencing the severity of traffic accidents based on the dataset of characteristic variables for autonomous and manual driving traffic accidents using machine learning; and constructing a joint prediction model for the severity of autonomous and manual driving accidents in a mixed traffic environment based on the characteristic variables influencing the severity of traffic accidents to predict the severity of accidents. The present invention improves prediction accuracy and the goodness of fit of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of accident cause analysis, and in particular to a method for predicting the severity of accidents in mixed autonomous driving and manual driving. Background Art

[0002] Before autonomous driving systems become fully widespread, mixed traffic conditions involving autonomous and manually driven vehicles will remain a typical feature of the roads. However, autonomous vehicles have also exposed numerous issues in real-world mixed traffic environments. Traffic safety analysis based on historical accident data, the construction of accident severity prediction models, quantitative research on the factors influencing accident severity in mixed traffic environments, and the evaluation of the safety effectiveness of autonomous vehicles will help develop targeted improvement strategies and establish an effective and comprehensive accident warning mechanism. This is crucial for improving the conditions under which autonomous vehicles can operate on the road, enhancing safety in mixed traffic environments, and reducing accident losses.

[0003] From a data perspective, the road traffic system is composed of four major elements: people, vehicles, roads, and the environment. Traffic accidents are often random events caused by malfunctions in these elements. The interaction of these elements determines the occurrence, development, and consequences of accidents. Therefore, when building accident severity prediction models in mixed traffic environments, it is necessary to use autonomous vehicles and manually driven vehicles as a starting point, study their behavioral changes during driving, trace their causes, and analyze factors such as road characteristics and the built environment that influence vehicle behavior to accurately predict accident severity. However, existing autonomous driving accident data is missing key variables, making it difficult to fully explore the factors affecting accident severity. In addition, existing datasets have small sample sizes and are subject to missing or unverified data, resulting in poor generalization of existing models and difficulty reflecting their performance in real-world scenarios.

[0004] At the methodological level, existing traffic safety analysis methods typically treat the severity of autonomous driving accidents and manually driven accidents as independent analysis units, using logit and probit regression models, as well as their extended models, to predict accidents and analyze the relationship between accident severity and various influencing factors. However, in mixed traffic environments, certain factors may affect the severity of accidents involving both autonomous vehicles and manually driven vehicles. Existing single models often ignore the mutual relationship between the severity of accidents between these two types of vehicles. In addition, the background factors of real-world accidents are complex and diverse, making it impossible to fully collect and quantify all factors that may affect accident severity, thereby introducing unobserved heterogeneity. Existing accident severity prediction methods in mixed environments ignore the mutual relationship between different types of vehicles and the unobserved heterogeneity in the dataset, resulting in biased model parameter estimates. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides a method for predicting the severity of accidents in mixed autonomous driving and manual driving.

[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0007] A method for predicting the severity of an accident in a mixed autonomous driving and manual driving situation comprises the following steps:

[0008] Obtain traffic accident data in mixed autonomous and manual driving environments;

[0009] According to the accident location in the traffic accident data, the road traffic characteristics and built environment characteristics of the accident location are extracted;

[0010] Match traffic accident data with road traffic characteristics and built environment characteristics at the accident location to establish a data set of traffic accident characteristic variables for both autonomous driving and manual driving.

[0011] Based on the characteristic variable dataset of traffic accidents involving autonomous driving and manual driving, we screen the characteristic variables that affect the severity of traffic accidents using machine learning.

[0012] Based on the characteristic variables that affect the severity of traffic accidents, a joint prediction model for the severity of accidents caused by autonomous driving and manual driving is constructed in a mixed traffic environment to predict the severity of accidents.

[0013] Furthermore, obtaining traffic accident data in a mixed environment of autonomous driving and manual driving includes:

[0014] Extract autonomous driving accident feature information from historical autonomous driving accident data and assign a unique accident code to each autonomous driving accident. Accident features include accident time, accident location, collision type, vehicle characteristics, natural environmental conditions, and a detailed description of the accident.

[0015] Locate the accident location based on the accident location and accident details, and convert it into latitude and longitude coordinates in the earth coordinate system;

[0016] Extracting characteristic information of manual driving accidents based on historical manual driving accident data;

[0017] Based on the characteristic information of autonomous driving accidents and manual driving accidents, all manual driving accidents with the same road characteristics and built environment characteristics as the autonomous driving accidents are screened within the set first buffer zone, and the manual driving accidents closest to the autonomous driving accidents are selected for pairing to generate traffic accident data in a mixed autonomous driving and manual driving environment.

[0018] Furthermore, based on the accident location in the traffic accident data, the road traffic characteristics and built environment characteristics of the accident location are extracted, including:

[0019] Extract road feature information at the accident location based on the accident location in the traffic accident data; road features include intersection type, traffic control device type, central divider type, number of lanes, crosswalks, and roadside parking spaces;

[0020] Extract the road attribute information of the accident location according to the accident location in the traffic accident data; the road attributes include road name, road grade, speed limit, number of lanes and one-way roads;

[0021] Based on the accident location in the traffic accident data, the built environment information of the accident location within the set second buffer zone is extracted; the built environment includes land use types, parks, restaurants, schools, bus and subway stations, hospitals and shopping centers.

[0022] Furthermore, the traffic accident data is matched with the road traffic characteristics and built environment characteristics of the accident location to establish a data set of traffic accident characteristic variables for autonomous driving and manual driving, including:

[0023] According to the accident codes and longitude and latitude coordinates, the characteristic information of autonomous driving accidents, the characteristic information of manual driving accidents, the road characteristic information, the road attribute information and the built environment information are matched to establish a characteristic variable dataset of autonomous driving and manual driving traffic accidents.

[0024] Furthermore, based on the characteristic variable dataset of traffic accidents involving autonomous driving and manual driving, we screened characteristic variables that influence the severity of traffic accidents based on machine learning, including:

[0025] Based on the characteristic variable dataset of autonomous driving and manual driving traffic accidents, a random forest model was constructed with the severity of autonomous driving accidents and manual driving accidents as the target variable, and accident characteristics, road characteristics, and built environment characteristics as feature variables.

[0026] In each decision tree of the random forest, the Gini index is used to measure the classification contribution of the feature variable to the accident severity when the node is split. The node sample is split into two subsets based on the feature variable, and the change in the node Gini index after the split is calculated.

[0027] Traverse all decision trees in the random forest, accumulate the changes in the Gini index caused by the feature variables in each node split, and obtain the feature variable importance of the feature variable in the random forest for the classification effect of accident severity;

[0028] The importance of all characteristic variables is standardized to obtain the relative importance of each characteristic variable to the severity of the accident;

[0029] The characteristic variables are sorted from high to low according to their relative importance, the cumulative contribution ratio is calculated based on the first number of characteristic variables, and the characteristic variables that affect the severity of the traffic accident are determined based on the cumulative contribution ratio of the characteristic variables and the set threshold.

[0030] Furthermore, the change in the Gini index of a node after splitting is calculated as follows:

[0031]

[0032] Among them, ΔGini(X j ) is the change in the node Gini index after splitting, Gini(D) is the Gini index of node sample D, |D L | is the child node dataset D after splitting L The size of |D R | is the child node dataset D after splitting R The size of |D| is the size of the current node dataset D, Gini(D L ) is the child node dataset D L The Gini index, Gini(D R ) is the child node dataset D R The Gini index, X j is the jth characteristic variable.

[0033] Furthermore, the cumulative contribution ratio of the characteristic variables is calculated as follows:

[0034] The first number of characteristic variables ranked in the top is accumulated, and the ratio of the sum of the accumulated characteristic variables to the sum of all characteristic variables is calculated to obtain the cumulative contribution ratio of the characteristic variables.

[0035] Furthermore, based on the characteristic variables affecting traffic accident severity, a joint prediction model for the severity of accidents involving autonomous driving and manual driving is constructed in a mixed traffic environment to predict accident severity, including:

[0036] A random parameter bivariate Probit model was constructed with the severity of autonomous driving accidents and manual driving accidents as the dependent variable and the characteristic variables affecting the severity of traffic accidents as the independent variable.

[0037] Assuming that the regression coefficient of the independent variable in the model varies randomly between different accident samples, set a new regression coefficient of the independent variable;

[0038] The joint probability density function of all accident data is established based on the independent variables and the regression coefficients of the new independent variables, and converted into a log-likelihood function;

[0039] The quasi-Monte Carlo method is used to approximate the integral in the log-likelihood function;

[0040] Based on the approximate log-likelihood function, the parameters are iteratively estimated using the gradient descent method;

[0041] Through multiple iterative optimizations until the log-likelihood function converges, the final coefficient value of each characteristic variable for the accident severity classification is obtained.

[0042] Furthermore, the quasi-Monte Carlo method is used to approximate the integral of the log-likelihood function:

[0043]

[0044] in, is the joint probability density function, Y i1 is the target random variable related to the classification result of the severity of the autonomous driving accident in the i-th sample, Y i2 is the target random variable related to the classification result of the severity of human driving accidents in the i-th sample, X i are the accident characteristics, road characteristics, and built environment variables related to the severity of the i-th sample, represents the random parameter related to the severity of the autonomous driving accident in the i-th sample, is the random parameter related to the severity of manual driving accidents in sample i, Σ is the covariance matrix between the severity of automatic driving and manual driving accidents, and M is the number of simulated sampling.

[0045] Furthermore, based on the approximate log-likelihood function, the gradient descent method is used to iteratively estimate the parameters, specifically:

[0046]

[0047] in, and are the random parameters β′ at the t+1th iteration i,1 The mean and variance of and are the random parameters β′ at the tth iteration i,1 The mean and variance of , η is the learning rate, is the gradient of the log-likelihood function.

[0048] The present invention has the following beneficial effects:

[0049] (1) The present invention obtains traffic accident data in a mixed environment of autonomous driving and manual driving from multiple data sources, matches the road characteristics and built environment characteristics of the accident location, and constructs a complete accident feature dataset to solve the problems of insufficient existing data volume and lack of key feature variables.

[0050] (2) This invention combines autonomous driving with manual driving in a mixed traffic environment to construct a joint accident severity prediction model. This model can analyze the impact of accident characteristics, road characteristics, and built environment characteristics on accident severity. It also improves prediction accuracy and model goodness of fit by overcoming the problems of existing models that ignore correlations between dependent variables and unobserved heterogeneity in the dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of a method for predicting the severity of an accident in mixed autonomous and manual driving situations is provided;

[0052] Figure 2 Extracting geographic coordinates of autonomous driving accidents;

[0053] Figure 3 Schematic diagram of pairing autonomous driving accidents with manual driving accidents;

[0054] Figure 4 Extracting process intent from the characteristic variable dataset of traffic accidents between autonomous driving and manual driving. DETAILED DESCRIPTION

[0055] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0056] like Figure 1 As shown, an embodiment of the present invention provides a method for predicting the severity of an accident in mixed autonomous driving and manual driving, comprising the following steps S1 to S5:

[0057] S1. Obtain traffic accident data in a mixed environment of autonomous driving and manual driving;

[0058] In an optional embodiment of the present invention, step S1 of acquiring traffic accident data in a mixed environment of automatic driving and manual driving includes:

[0059] Extract autonomous driving accident feature information from historical autonomous driving accident data and assign a unique accident code to each autonomous driving accident. Accident features include accident time, accident location, collision type, vehicle characteristics, natural environmental conditions, and a detailed description of the accident.

[0060] Locate the accident location based on the accident location and accident details, and convert it into latitude and longitude coordinates in the earth coordinate system;

[0061] Extracting characteristic information of manual driving accidents based on historical manual driving accident data;

[0062] Based on the characteristic information of autonomous driving accidents and manual driving accidents, all manual driving accidents with the same road characteristics and built environment characteristics as the autonomous driving accidents are screened within the set first buffer zone, and the manual driving accidents closest to the autonomous driving accidents are selected for pairing to generate traffic accident data in a mixed autonomous driving and manual driving environment.

[0063] In this embodiment, step S1 collects data on 699 autonomous driving accidents that occurred from January 2015 to July 2024 from the National Highway Traffic Safety Administration, the California Department of Motor Vehicles, global news reports, and social media posts, and collects factors such as accident severity, collision type, and accident time.

[0064] This embodiment uses the historical autonomous driving accident data published by the National Highway Traffic Safety Administration (NHTSA) and the California Department of Motor Vehicles (CA DMV) as the main source to extract autonomous driving accident feature information and set an independent accident code for each autonomous driving accident. Global news reports and social media posts are used as auxiliary data sources, and large language models (such as ChatGPT) are used to uniformly convert accident information from different countries into English. In order to ensure the accuracy and consistency of the data, specific prompt questions are set in ChatGPT (for example: estimating severity, extracting accident type, extracting accident time, etc.) to extract accident feature information from text data. Accident features include: accident time, location, collision type, vehicle characteristics, natural environmental conditions, and detailed description of the accident.

[0065] In this embodiment, step S1 uses the accident location and accident details to locate the specific location of the accident in Google Earth and converts it into longitude and latitude coordinates in the WGS1984 coordinate system. Figure 2 shown.

[0066] In this embodiment, step S1 collects manual driving accident data from January 2015 to July 2024 from the traffic injury mapping system of the University of California, Berkeley, and obtains factors such as accident severity, collision type, accident time, and accident coordinates; the two types of accidents are mapped to ArcGIS, and a 50-meter buffer zone is constructed to match each autonomous driving accident with a manual driving accident with the same road and built environment characteristics. Specifically, Figure 3 shown.

[0067] This example uses the University of California, Berkeley's Traffic Injury Mapping System as the primary data source to extract historical human-driven accident data, including each accident's characteristics and geographic coordinates. Both autonomous driving and human-driven accidents are imported into a geographic information system (ArcGIS). For each autonomous driving accident, a 50m-radius buffer zone is set to filter out all human-driven accidents that share the same road and built environment characteristics as the autonomous driving accident. Based on the distance from the autonomous driving accident site, the nearest human-driven accident within the buffer zone is selected for pairing to generate traffic accident data for mixed traffic environments.

[0068] S2. Extracting road traffic characteristics and built environment characteristics of the accident location based on the accident location in the traffic accident data;

[0069] In an optional embodiment of the present invention, step S2 extracts road traffic characteristics and built environment characteristics of the accident location based on the accident location in the traffic accident data, including:

[0070] Extract road feature information at the accident location based on the accident location in the traffic accident data; road features include intersection type, traffic control device type, central divider type, number of lanes, crosswalks, and roadside parking spaces;

[0071] Extract the road attribute information of the accident location according to the accident location in the traffic accident data; the road attributes include road name, road grade, speed limit, number of lanes and one-way roads;

[0072] Based on the accident location in the traffic accident data, the built environment information of the accident location within the set second buffer zone is extracted; the built environment includes land use types, parks, restaurants, schools, bus and subway stations, hospitals and shopping centers.

[0073] In this embodiment, if Figure 4 As shown, step S2 first uses the longitude and latitude of the accident to collect road feature information such as intersection type, traffic control facility type, and number of lanes on Google Earth; uses Overpass Turb to query road attribute information such as road grade, road speed limit, and one-way road at the accident site; uses ArcGIS to collect land use type and built environment information such as the presence of parks, schools, restaurants, subway stations, and bus stops at the accident site; thereby matching the road traffic characteristics and built environment characteristics of the accident location.

[0074] This example uses Google Earth Street View to locate the accident location based on the accident's latitude and longitude coordinates and extract relevant road feature information. This includes, but is not limited to, location, intersection type, traffic control device type, median type, number of lanes, crosswalks, and roadside parking spaces.

[0075] This example uses the Overpass API to query road information near the accident site using the Overpass Turb tool. Through API requests, the Overpass Turb platform will return road attribute information such as road name, road grade, speed limit, number of lanes, and one-way roads within the specified latitude and longitude.

[0076] This example uses a Python script to call the Overpass API, sending an HTTP request and setting query parameters to specify the query area and the built environment information to be queried. After the API returns data in JSON format, the response data is parsed into a Python dictionary object using the response.Json() parameter. The built environment elements in the dictionary object are traversed, and the label information and latitude and longitude coordinates of each element in the WGS1984 coordinate system are extracted. The autonomous driving accident and built environment data are imported into ArcGIS. A buffer with a radius of 500m is set at each autonomous driving accident point. The spatial join tool is used to extract the built environment characteristics within the buffer area, including land use types, parks, restaurants, schools, bus and subway stations, hospitals, shopping malls, etc.

[0077] S3. Match traffic accident data with road traffic characteristics and built environment characteristics at the accident location to establish a data set of traffic accident characteristic variables for autonomous driving and manual driving.

[0078] In an optional embodiment of the present invention, step S3 matches the traffic accident data with the road traffic characteristics and built environment characteristics of the accident location to establish a data set of traffic accident characteristic variables for autonomous driving and manual driving, including:

[0079] According to the accident codes and longitude and latitude coordinates, the characteristic information of autonomous driving accidents, the characteristic information of manual driving accidents, the road characteristic information, the road attribute information and the built environment information are matched to establish a characteristic variable dataset of autonomous driving and manual driving traffic accidents.

[0080] In this embodiment, step S3 uses the accident code and longitude and latitude as common identifiers to match the autonomous driving accident feature information, manual driving accident feature information, road feature information, road attribute information, and built environment features. The matched information is regarded as a complete accident feature variable sample, and independent coding is used to convert character variables into numerical variables to establish the accident feature variable dataset D2.

[0081] This embodiment uses two common identifiers, accident codes and longitude and latitude, to match autonomous driving accident feature information, manual driving accident feature information, road feature information, road attribute information, and built environment information. This matched information is considered a complete sample of accident feature variables. For all accident samples, the integrated feature variables form the feature variable dataset D1. During the matching process, accident codes and longitude and latitude coordinates are used for precise alignment, ensuring that each accident data point is correctly associated with the relevant road and built environment features.

[0082] This embodiment preprocesses the sample data set D1, first using the mode to fill in the missing categorical variable values. For outliers or data that do not meet the predefined standards (for example, unreasonable accident time, geographical location), they are regarded as "dirty data" and deleted. Then, the independent coding method is used to convert all n categorical variables in the sample data set into n-1 independent binary variables to avoid the virtual variable trap. Among them, the ignored variables are reference variables, which represent the most common or most meaningful categories in practical applications. Finally, the characteristic variable dataset D2 of autonomous driving and manual driving traffic accidents is established, and the results are shown in Table 1:

[0083] Table 1 shows the accident feature variable dataset after variable extraction and feature coding.

[0084]

[0085] S4. Based on the dataset of characteristic variables of traffic accidents involving autonomous driving and manual driving, filter characteristic variables that influence the severity of traffic accidents using machine learning;

[0086] In an optional embodiment of the present invention, step S4 screens the traffic accident severity-influencing characteristic variables based on machine learning according to the autonomous driving and manual driving traffic accident characteristic variable datasets, including:

[0087] Based on the characteristic variable dataset of autonomous driving and manual driving traffic accidents, a random forest model was constructed with the severity of autonomous driving accidents and manual driving accidents as the target variable, and accident characteristics, road characteristics, and built environment characteristics as feature variables.

[0088] In each decision tree of the random forest, the Gini index is used to measure the classification contribution of the feature variable to the accident severity when the node is split. The node sample is split into two subsets based on the feature variable, and the change in the node Gini index after the split is calculated.

[0089] Traverse all decision trees in the random forest, accumulate the changes in the Gini index caused by the feature variables in each node split, and obtain the feature variable importance of the feature variable in the random forest for the classification effect of accident severity;

[0090] The importance of all characteristic variables is standardized to obtain the relative importance of each characteristic variable to the severity of the accident;

[0091] The characteristic variables are sorted from high to low according to their relative importance, the cumulative contribution ratio is calculated based on the first number of characteristic variables, and the characteristic variables that affect the severity of the traffic accident are determined based on the cumulative contribution ratio of the characteristic variables and the set threshold.

[0092] The calculation method of the cumulative contribution ratio of the characteristic variable is:

[0093] The first number of characteristic variables ranked in the top is accumulated, and the ratio of the sum of the accumulated characteristic variables to the sum of all characteristic variables is calculated to obtain the cumulative contribution ratio of the characteristic variables.

[0094] In this embodiment, step S4 constructs a random forest model based on the sample dataset D2 and selects the top 20 variables with the highest eigenvector importance. These variables include rear-end collision, side collision, side-scrubbing, night, slippery road surface, bus lane, traffic light, two-way three-lane road, bus stop and subway station, weekday, pedestrian crossing, bicycle lane, school, park, two-way four-lane road, and Y-shaped intersection.

[0095] This embodiment constructs a random forest model based on the sample data set D2, using the severity of autonomous driving accidents and manual driving accidents as target variables, and accident characteristics, road characteristics, and built environment characteristics as feature variables.

[0096] In this embodiment, the Gini index is used to measure the feature variable X in each decision tree of the random forest. j The classification contribution to the severity of the accident when the node splits is calculated as follows:

[0097]

[0098] Where c is the number of accident severity categories, p i It is the sample proportion of a certain type of accident in the node (such as the proportion of accidents without injury).

[0099] Through the characteristic variable X j Split the node sample D into two subsets D L and D R , and calculate the change in the node Gini index after the split, which reflects the characteristic variable X j The influence degree on the classification of accident severity. The larger the value of the change, the greater the contribution of the characteristic variable to the classification of accident severity. The calculation formula is:

[0100]

[0101] Among them, ΔGini(X j ) is the change in the node Gini index after splitting, Gini(D) is the Gini index of node sample D, |D L | is the child node dataset D after splitting L The size of |D R | is the child node dataset D after splitting R The size of |D| is the size of the current node dataset D, Gini(D L ) is the child node dataset D L The Gini index, Gini(D R ) is the child node dataset D R The Gini index, X j is the jth characteristic variable.

[0102] This example traverses all decision trees in the random forest and accumulates the feature variables X j The change in the Gini index caused by each node split is the characteristic variable X j The overall impact of the accident severity classification in the random forest, that is, the importance of the feature variable. The calculation formula is:

[0103]

[0104] Among them, T is the set of all trees in the random forest, t is the serial number of each tree, and i is the serial number of each node in the tree.

[0105] In this embodiment, the importance of all characteristic variables is standardized to obtain the relative importance of each characteristic variable to the severity of the accident. The calculation formula is:

[0106]

[0107] Where n is the total number of feature variables in the model.

[0108] This embodiment sorts the feature variables from high to low according to the standardized feature variable importance. For the top k feature variables, their importance scores are accumulated and calculated, and the comprehensive score is divided by the total importance score of all feature variables to obtain the cumulative contribution ratio. The calculation formula is:

[0109]

[0110] According to the set 80% threshold, the selected feature variables are determined.

[0111] S5. Based on the characteristic variables affecting the severity of traffic accidents, a joint prediction model for the severity of accidents involving autonomous driving and manual driving is constructed in a mixed traffic environment to predict the severity of accidents.

[0112] In an optional embodiment of the present invention, step S5 constructs a joint prediction model for the severity of accidents caused by autonomous driving and manual driving in a mixed traffic environment based on characteristic variables affecting the severity of traffic accidents to predict the severity of accidents, including:

[0113] A random parameter bivariate Probit model was constructed with the severity of autonomous driving accidents and manual driving accidents as the dependent variable and the characteristic variables affecting the severity of traffic accidents as the independent variable.

[0114] Assuming that the regression coefficient of the independent variable in the model varies randomly between different accident samples, set a new regression coefficient of the independent variable;

[0115] The joint probability density function of all accident data is established based on the independent variables and the regression coefficients of the new independent variables, and converted into a log-likelihood function;

[0116] The quasi-Monte Carlo method is used to approximate the integral in the log-likelihood function;

[0117] Based on the approximate log-likelihood function, the parameters are iteratively estimated using the gradient descent method;

[0118] Through multiple iterative optimizations until the log-likelihood function converges, the final coefficient value of each characteristic variable for the accident severity classification is obtained.

[0119] In this embodiment, step S5 establishes a random parameter bivariate Probit with the severity of autonomous driving accidents and manual driving accidents as dependent variables and the top 20 eigenvector importance-ranked eigenvalues ​​as independent variables to obtain the parameter estimation results of the final model.

[0120] This example is based on the sample data set D2, and uses the severity of autonomous driving accidents and manual driving accidents as dependent variables, and the filtered characteristic variables as independent variables to construct a bivariate Probit model. Specifically,

[0121] if y i,1 =0otherwise

[0122] if y i,2 =0 otherwise

[0123] Among them, y i,1 and y i,2 Represent the severity results of autonomous driving accidents and manual driving accidents, respectively, and and Denotes the corresponding latent variable. X i,1 and Xi,2 β is the independent variable, representing the accident characteristics, road characteristics and built environment variables related to the severity of the accident. i,1 and β i,2 is the regression coefficient of the corresponding independent variable, ε i,1 and ε i,2 are the errors, which follow a standard normal distribution with mean 0 and variance 1.

[0124] Considering that there is a correlation between the severity of autonomous driving accidents and manual driving accidents, the error term ε i,1 and ε i,2 The correlation between them is expressed by the covariance matrix, specifically:

[0125]

[0126] Where ρ is the correlation coefficient, which represents the correlation measure of the severity of autonomous driving and manual driving accidents.

[0127] This example sets random parameters and assumes that the regression coefficients of the independent variables in the model vary randomly between different accident samples to capture the heterogeneity between individuals. Specifically,

[0128] β′ i,1 =β i,1 +δ i,1

[0129] β′ i,2 =β i,2 +δ i,2

[0130] Among them, δ i,1 and δ i,2 represents a randomly distributed term that captures unobserved heterogeneity in the accident data.

[0131] In this embodiment, for the i-th accident, P(Y i1 ,Y i2 ∣X i ,β′ i,1 ,β′ i,2 ,Σ) is the given independent variable X i and parameter β′ i,1 ,β′ i,2 , joint probability density under Σ conditions. Assuming that the N accident data are independently distributed, the likelihood function Represents the joint probability density function of all accident data, specifically:

[0132]

[0133] Where Σ is the covariance matrix between the severity of autonomous driving and manual driving accidents, and ε i,1 and εi,2 The correlation of f(β′ i,1 ,β′ i,2 ) is the random parameter β′ i,1 ,β′ i,2 The joint distribution of .

[0134] Since the model contains random parameters, it needs to be converted into a log-likelihood function. The calculation formula is:

[0135]

[0136] This embodiment uses the quasi-Monte Carlo method to approximate the integral in the log-likelihood function, setting ε i,1 and ε i,2 For the random effect obtained by sampling M times from the normal distribution, the integral in the likelihood function is approximated by the formula:

[0137]

[0138] Where, M is the number of simulated samples; represents the random effect parameter related to the severity of autonomous driving accidents in the i-th sample, is the random effect parameter related to the severity of manual driving accidents in sample i, which are all derived from the joint distribution f(β′ i,1 ,β′ i,2 ) is sampled m times.

[0139] For each accident sample i, its log-likelihood function is approximated as:

[0140]

[0141] in, is the joint probability density function; Y i1 is the target random variable of the i-th sample, reflecting the classification result of the severity of the autonomous driving accident; Y i2 is the target random variable of the i-th sample, reflecting the classification result of the severity of the manual driving accident; X i are the accident characteristics, road characteristics, and built environment variables related to the severity of accident i; Σ is the covariance matrix between the severity of autonomous driving and manual driving accidents; M is the number of simulation samples. This embodiment uses the gradient descent method to iteratively estimate parameters based on the approximate log-likelihood function. Assume that the random parameter β′ i,1 and β′ i,2 Normal distribution:

[0142]

[0143] Calculate the gradient of the distribution parameters, the calculation formula is (with random parameter β′i,1 For example):

[0144]

[0145] in, is from the distribution f(β′ i,1 ,β′ i,2 ) draws the mth random parameter sample drawn from .

[0146] Update the independent variable regression coefficient according to the gradient of the distribution function, specifically:

[0147]

[0148]

[0149] in, and are the random parameters β′ at the t+1th iteration i,1 The mean and variance of and are the random parameters β′ at the tth iteration i,1 The mean and variance of , η is the learning rate, is the gradient of the log-likelihood function.

[0150] Through multiple iterations of optimization until the log-likelihood function converges, the final coefficient value for each characteristic variable for the accident severity classification is obtained. If the mean and variance of a random parameter are not significant at the 0.1 significance level, it is set as a fixed parameter and the estimation process is repeated until all parameters are significant. The parameter estimation results of the final model are shown in Table 2:

[0151] Table 2 shows the parameter estimation results of the random parameter bivariate Probit model

[0152]

[0153]

[0154] Based on the above parameter estimation results, a method for predicting the target variable y is constructed. i1 and y i2 The random parameter bivariate probit model is:

[0155] if y i,1 =0 otherwise

[0156] if y i,2 =0 otherwise

[0157] where X i1 and X i2 is the characteristic variable determined in the parameter estimation results; β′ i,1 ,β′ i,2 are random effects, all with mean μ and variance σ 2 The coefficients in the table represent the random parameters μ and σ 2 .

[0158] Calculate the joint probability of the target variable and predict the severity of autonomous driving accidents and manual driving accidents as follows:

[0159]

[0160] where Φ2 is the cumulative distribution function of the bivariate normal distribution, Represents the characteristic variable X i A linear combination of the model parameters (mean vector μ1, μ2).

[0161] When P(Y i1 =1,Y i2 =1|X i )>0.5, then predict y i1 =1,y i2 =1, indicating that both autonomous driving accidents and manual driving accidents result in serious injuries. If the above conditions are not met, it is predicted that at least one of the accident types does not result in serious injuries.

[0162] In an optional embodiment of the present invention, in order to compare and analyze the effectiveness of the method of the present invention, a random parameter univariate model (RUB) and a bivariate Probit (BP) model are simultaneously established, and the AIC value and likelihood ratio test are used to compare with the method of the present invention.

[0163] This example is based on the Akaike Information Criterion (AIC), taking into account the complexity and fitting effect of the model, and calculates the relative goodness of fit of the model. The calculation formula is:

[0164] AIC=-2 ln(L)+2k

[0165] Where k is the number of parameters in the model and L is the maximum likelihood estimate of the model.

[0166] This example compares the likelihood function values ​​of the two models based on the likelihood ratio test (LC) to evaluate whether there is a significant difference in the fitting effect between the models. The calculation formula is:

[0167] LR=-2 ln(L N )-ln(L A )

[0168] Among them, ln(L N ) and ln(L A ) represent the log-likelihood values ​​of model N and model A respectively.

[0169] When the P value corresponding to LR is less than 0.05, model A is considered to be significantly better than model N.

[0170] This example uses the AUC value, accuracy, precision, and recall rate to evaluate the prediction accuracy of the model based on the actual accident severity categories of the accident samples and the accident severity categories predicted by the model. The calculation formula is:

[0171]

[0172] Among them, TP is the number of samples whose actual severity is injury accidents and the model also predicts them as injury accidents; FP is the number of samples whose actual severity is no injury accidents but the model incorrectly predicts them as injury accidents; TN is the number of samples whose actual severity is no injury accidents and the model also predicts them as no injury accidents; FN is the number of samples whose actual severity is injury accidents but the model incorrectly predicts them as no injury accidents.

[0173] The comparison results of the method of the present invention with the random parameter univariate model (RUB) and the bivariate Probit (BP) model are shown in Table 3:

[0174] Table 3 Goodness of fit of the three models

[0175]

[0176] Based on the goodness of fit test results of the above three models, the AIC values ​​of the random parameter univariate Probit model, the bivariate Probit model and the random parameter bivariate Probit model are 1364.1, 1343.1 and 1334.7 respectively. The AIC value of the random parameter bivariate Probit model is significantly reduced (the difference is greater than 5 compared with other models), indicating that the random parameter bivariate Probit model provides an excellent fitting effect. In addition, the likelihood ratio test results confirm that the random parameter bivariate Probit model is statistically significantly better than the random parameter univariate Probit model and the bivariate. Therefore, the random parameter bivariate Probit model proposed in the present invention can significantly improve the goodness of fit of the accident severity in a mixed environment and has certain practical value.

[0177] Furthermore, the model's prediction accuracy was evaluated based on the AUC, accuracy, precision, and recall. The model achieved an AUC of 0.74, a prediction accuracy of 0.66, a prediction precision of 0.65, and a prediction recall of 0.69. The values ​​of these evaluation metrics were all significantly greater than 0.5, demonstrating the excellent performance of the method provided by this invention, with high accuracy in predicting accident severity in mixed autonomous and human-driven environments.

[0178] In summary, this paper begins with raw autonomous driving accident data, matches it with nearby human-driven accidents, and uses multiple methods to capture vehicle motion, road characteristics, and built environment features, thereby constructing a dataset of traffic accident characteristic variables in mixed traffic environments. By filtering out redundant attributes through correlation analysis and independently encoding character data, the original accident dataset is transformed into a dataset amenable to statistical mining and analysis. Furthermore, a random parameter bivariate probit model is introduced to address the issues of existing accident severity prediction models that ignore correlations between dependent variables and unobserved heterogeneity in the dataset, thereby improving the model's prediction accuracy and goodness of fit.

[0179] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0180] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1A step that specifies a function in one or more boxes.

[0182] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

[0183] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A method for predicting the severity of accidents in mixed autonomous and manual driving situations, characterized in that: The following steps are involved: Obtain traffic accident data in mixed autonomous and manual driving environments, including: Extract autonomous driving accident feature information from historical autonomous driving accident data and assign a unique accident code to each autonomous driving accident. Accident features include accident time, accident location, collision type, vehicle characteristics, natural environmental conditions, and a detailed description of the accident. Locate the accident location based on the accident location and accident details, and convert it into latitude and longitude coordinates in the earth coordinate system; Extracting characteristic information of manual driving accidents based on historical manual driving accident data; Based on the characteristic information of autonomous driving accidents and manual driving accidents, all manual driving accidents with the same road characteristics and built environment characteristics as the autonomous driving accidents are screened within the set first buffer zone. The manual driving accidents closest to the autonomous driving accidents are selected for pairing, thereby generating traffic accident data for a mixed autonomous driving and manual driving environment. According to the accident location in the traffic accident data, the road traffic characteristics and built environment characteristics of the accident location are extracted, including: Extract road feature information at the accident location based on the accident location in the traffic accident data; road features include intersection type, traffic control device type, central divider type, number of lanes, crosswalks, and roadside parking spaces; Extract the road attribute information of the accident location according to the accident location in the traffic accident data; the road attributes include road name, road grade, speed limit, number of lanes and one-way roads; Based on the accident location in the traffic accident data, the built environment information within the set second buffer zone of the accident location is extracted; the built environment includes land use types, parks, restaurants, schools, bus and subway stations, hospitals, and shopping centers; Traffic accident data is matched with the road traffic characteristics and built environment characteristics of the accident location points to establish a dataset of traffic accident characteristic variables for autonomous driving and manual driving, including: Based on the accident codes and longitude and latitude coordinates, the characteristic information of autonomous driving accidents, the characteristic information of manual driving accidents, road characteristics, road attribute information, and built environment information are matched to establish a characteristic variable dataset of autonomous driving and manual driving traffic accidents. Based on the characteristic variable dataset of traffic accidents involving autonomous driving and manual driving, we screen the characteristic variables that affect the severity of traffic accidents using machine learning. Based on the characteristic variables that influence traffic accident severity, a joint prediction model for the severity of autonomous driving and manual driving accidents in mixed traffic environments is constructed to predict accident severity, including: A random parameter bivariate Probit model was constructed with the severity of autonomous driving accidents and manual driving accidents as the dependent variable and the characteristic variables affecting the severity of traffic accidents as the independent variable. Assuming that the regression coefficient of the independent variable in the model varies randomly between different accident samples, set a new regression coefficient of the independent variable; The joint probability density function of all accident data is established based on the independent variables and the regression coefficients of the new independent variables, and converted into a log-likelihood function; The quasi-Monte Carlo method is used to approximate the integral in the log-likelihood function; Based on the approximate log-likelihood function, the parameters are iteratively estimated using the gradient descent method; Through multiple iterative optimizations until the log-likelihood function converges, the final coefficient value of each characteristic variable for the accident severity classification is obtained.

2. The method for predicting the severity of accidents in mixed autonomous and manual driving situations according to claim 1, characterized in that: Based on the characteristic variable dataset of traffic accidents involving autonomous driving and manual driving, we screen the characteristic variables that affect the severity of traffic accidents based on machine learning, including: Based on the characteristic variable dataset of autonomous driving and manual driving traffic accidents, a random forest model was constructed with the severity of autonomous driving accidents and manual driving accidents as the target variable, and accident characteristics, road characteristics, and built environment characteristics as feature variables. In each decision tree of the random forest, the Gini index is used to measure the classification contribution of the feature variable to the accident severity when the node is split. The node sample is split into two subsets based on the feature variable, and the change in the node Gini index after the split is calculated. Traverse all decision trees in the random forest, accumulate the changes in the Gini index caused by the feature variables in each node split, and obtain the feature variable importance of the feature variable in the random forest for the classification effect of accident severity; The importance of all characteristic variables is standardized to obtain the relative importance of each characteristic variable to the severity of the accident; The characteristic variables are sorted from high to low according to their relative importance, the cumulative contribution ratio is calculated based on the first number of characteristic variables, and the characteristic variables that affect the severity of the traffic accident are determined based on the cumulative contribution ratio of the characteristic variables and the set threshold.

3. The method for predicting the severity of accidents in mixed automatic and manual driving situations according to claim 2, characterized in that: The calculation method for the change in the node Gini index after splitting is: Among them, ΔGini(X j ) is the change in the node Gini index after splitting, Gini(D) is the Gini index of node sample D, |D L | is the child node dataset D after splitting L The size of |D R | is the child node dataset D after splitting R The size of |D| is the size of the current node dataset D, Gini(D L ) is the child node dataset D L The Gini index, Gini(D R ) is the child node dataset D R The Gini index, X j is the jth characteristic variable.

4. The method for predicting the severity of accidents in mixed automatic and manual driving situations according to claim 2, characterized in that: The calculation method of the cumulative contribution ratio of the characteristic variable is: The first number of characteristic variables ranked in the top is accumulated, and the ratio of the sum of the accumulated characteristic variables to the sum of all characteristic variables is calculated to obtain the cumulative contribution ratio of the characteristic variables.

5. The method for predicting the severity of accidents in mixed automatic and manual driving situations according to claim 1, characterized in that: The quasi-Monte Carlo method is used to approximate the integral in the log-likelihood function: in, is the joint probability density function, Y i1 is the target random variable related to the classification result of the severity of the autonomous driving accident in the i-th sample, Y i2 is the target random variable related to the classification result of the severity of human driving accidents in the i-th sample, X i are the accident characteristics, road characteristics, and built environment variables related to the severity of the i-th sample, represents the random parameter related to the severity of the autonomous driving accident in the i-th sample, is the random parameter related to the severity of manual driving accidents in sample i, Σ is the covariance matrix between the severity of automatic driving and manual driving accidents, and M is the number of simulated sampling.

6. The method for predicting the severity of accidents in mixed automatic and manual driving situations according to claim 1, characterized in that: Based on the approximate log-likelihood function, the gradient descent method is used to iteratively estimate the parameters, specifically: in, and are the random parameters β′ at the t+1th iteration i,1 The mean and variance of and are the random parameters β′ at the tth iteration i,1 The mean and variance of , η is the learning rate, is the gradient of the log-likelihood function.

Citation Information

Patent Citations

  • Road traffic accident risk factor analysis method based on random forest

    CN110276370A

  • Road traffic accident severity and duration prediction method

    CN117010559A