A road safety analysis method and system for eliminating traffic accident sample heterogeneity
By constructing a clustering model and a safety analysis model, the safety influencing factors of each cluster are identified, solving the problem that existing technologies cannot consider the heterogeneity of accidents, and achieving a more accurate road safety assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2022-12-30
- Publication Date
- 2026-05-26
AI Technical Summary
Existing safety accident analysis models cannot effectively consider the heterogeneous impact of the accident itself and the heterogeneous effects between accident sample characteristics.
By constructing a clustering model, traffic accidents are divided into a predetermined number of clusters, and a safety analysis model is built based on each cluster. The analysis is conducted using a negative binomial regression model with random parameters to identify the safety influencing factors of each cluster.
It provides an accurate, comprehensive, and objective road safety evaluation method that can better reflect the authenticity of influencing data, has a wider range of applications, and can better characterize the impact of the same influencing factors on different accident groups.
Smart Images

Figure CN116384793B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of traffic safety technology, specifically relating to a road safety analysis method and system for eliminating the heterogeneity of traffic accident samples. Background Technology
[0002] In recent years, constructing road safety accident analysis models has become a research hotspot in the field of traffic safety. By building safety analysis models, the impact of various safety factors on accident occurrence can be explored. Currently, in both scientific research and patent applications, most studies introduce safety influencing factors and construct safety analysis models based on counting models (such as Poisson or negative binomial regression). In reality, the same influencing factors have different effects on different accident types. Therefore, many studies construct stochastic safety models to assess the heterogeneous impact of various safety factors. However, most existing studies, even with the construction of stochastic safety models, still cannot consider the heterogeneous impact of the accidents themselves, nor can they simultaneously consider the heterogeneous effects between factors and accident sample characteristics. Summary of the Invention
[0003] The technical problem to be solved by this invention is that most existing safety accident analysis model studies only construct stochastic safety models and still cannot consider the heterogeneous impact of the accident itself, nor can they simultaneously consider the heterogeneous effects between factors and accident sample characteristics.
[0004] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0005] A road safety analysis method for eliminating heterogeneity in traffic accident samples involves the following steps for evaluating the safety of roads in the study area:
[0006] Step A: For each traffic accident in the study area, based on the preset accident characteristics of each type of traffic accident and the preset road characteristics of each type of road in the study area, construct a clustering model to divide each traffic accident into a preset number of clusters.
[0007] Step B: Based on the traffic accidents corresponding to each ethnic group, and combined with the pre-defined road characteristics of each type in the study area, construct a safety analysis model corresponding to each ethnic group.
[0008] Step C: Based on the safety analysis models corresponding to each ethnic group, obtain the characteristics of various types of roads that affect road safety for each ethnic group, and realize the safety evaluation of roads in the study area.
[0009] As a preferred embodiment of the present invention, step A specifically involves the following steps: constructing a clustering model to divide each traffic accident into a preset number of clusters.
[0010] Step A1: For each traffic accident in the study area, based on the preset accident characteristics and road characteristics corresponding to each traffic accident, and combined with the preset number of clusters, construct the initial clustering model as shown below:
[0011]
[0012] In the formula, p(c|x i x represents the probability that traffic accident i belongs to group c, where c∈C; i This represents the preset accident characteristics corresponding to traffic accident i and the preset road characteristics corresponding to the study area; C is the total number of populations; α s α represents the regression parameter corresponding to population s. s This represents the constant corresponding to the population s;
[0013] Step A2: Based on the initial clustering model and a preset probability threshold, classify each traffic accident into a different cluster.
[0014] As a preferred embodiment of the present invention, the number of groups of the preset number of groups is obtained by iteratively performing the following steps:
[0015] Step 1: For each traffic accident in the study area, based on the preset accident characteristics and road characteristics corresponding to each traffic accident, and combined with the number of clusters (the initial number of clusters is 1), construct a clustering model.
[0016] Step 2: Based on the constructed population clustering model, obtain AIC and BIC;
[0017]
[0018]
[0019] In the formula, SSE represents the sum of squared errors for each cluster in the clustering model; I represents the total number of traffic accidents in the study area; L represents the total number of preset accident features and preset road features for each type.
[0020] Step 3: Based on AIC and BIC, and combined with AIC and BIC from the previous iteration, if the percentage change in AIC and BIC is less than the preset percentage threshold, then the number of populations corresponding to the current iteration is the preset number of populations, and the iteration ends; otherwise, increment the number of populations by one and return to Step 1.
[0021] As a preferred embodiment of the present invention, the security analysis models corresponding to each group in step B are obtained through the following process:
[0022] Step B1: Based on the traffic accidents corresponding to each ethnic group, and combined with the pre-defined road characteristics of each type in the study area, a safety analysis model corresponding to each ethnic group is constructed using a negative binomial regression model with random parameters. The model is shown in the following formula.
[0023]
[0024] In the formula; N ci X represents the total number of traffic accidents in group c; j Let ε represent the feature data of the j-th type of road in the group, ε be the regression model error corresponding to the group, β0 represent the constant corresponding to the group, and β be the feature data of the j-th type of road in the group. j is the regression parameter corresponding to the j-th type of road feature in the population; J represents the preset total number of road features of each type.
[0025] Step B2: For the security analysis model, based on β j The security analysis model has been revised and updated. The updated security analysis model is shown below:
[0026]
[0027] In the formula, It follows a normal distribution with a mean of 1 and a variance of δ. 2 .
[0028] As a preferred embodiment of the present invention, in step C, based on the safety analysis models corresponding to each ethnic group, the following process is specifically used to obtain the characteristics of various types of roads affecting the road safety of the target area corresponding to each ethnic group, thereby achieving a safety evaluation of the roads in the study area:
[0029] When β′ j If β' > 0, it indicates that the characteristics of road type j have a positive impact on the occurrence of accidents in the population. j If the value is less than 0, it indicates that the characteristics of road type j have a negative impact on the occurrence of accidents in the group.
[0030] As a preferred technical solution of the present invention, the preset accident characteristics include the age and gender of the accident participants, the weather conditions at the accident site, the signal control at the accident site, and whether the accident site is an intersection.
[0031] As a preferred technical solution of the present invention, the preset road characteristics of each type include road network density, road land area, green land area, main road density, and secondary road density.
[0032] A system for road safety analysis based on eliminating heterogeneity of traffic accident samples includes a clustering module, a safety analysis module, and a safety evaluation module. The clustering module constructs a clustering model for each traffic accident in the study area based on the preset accident characteristics of each type of traffic accident and the preset road characteristics of each type of road in the study area, and divides each traffic accident into a preset number of clusters.
[0033] The safety analysis module constructs a safety analysis model for each group based on the traffic accidents corresponding to each group and the pre-defined road characteristics of each type in the study area.
[0034] The safety evaluation module obtains the road characteristics that affect road safety for each type of road based on the safety analysis model corresponding to each ethnic group.
[0035] A terminal for a road safety analysis method to eliminate heterogeneity in traffic accident samples includes a memory and a processor, which are communicatively connected. The memory stores computer instructions, and the processor executes the computer instructions to perform the road safety analysis method for eliminating heterogeneity in traffic accident samples.
[0036] A computer-readable medium for storing software includes instructions executable by one or more computers, which, when executed by the one or more computers, perform a road safety analysis method for eliminating heterogeneity in traffic accident samples.
[0037] The beneficial effects of this invention are as follows: This invention provides a road safety analysis method and system for eliminating the heterogeneity of traffic accident samples. First, accidents are clustered according to their occurrence characteristics, into groups with different attribute characteristics. Then, based on the safety influencing factors of the accidents, traffic safety models are performed on different groups. A safety analysis model is applied to obtain the influencing factors of traffic road safety in the affected area, and a safety evaluation of the area is conducted. Through the technical solution of this invention, a precise, comprehensive, objective, and truthful road safety evaluation method that reflects the impact data is provided, with a wider range of applications. In particular, clustering based on accident characteristics can amplify the heterogeneity between accident samples, better characterizing the impact of the same influencing factors on different accident groups. Attached Figure Description
[0038] Figure 1 This is a flowchart of the method in this embodiment. Detailed Implementation
[0039] The present invention will be further described below with reference to the accompanying drawings. The following embodiments will enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way.
[0040] like Figure 1 As shown, a road safety analysis method for eliminating heterogeneity in traffic accident samples is described. For the study area, the following steps are performed to achieve a safety evaluation of the roads in the study area:
[0041] Step A: For each traffic accident in the study area, based on the preset accident characteristics of each type of traffic accident and the preset road characteristics of each type of road in the study area, construct a clustering model to divide each traffic accident into a preset number of clusters.
[0042] In this embodiment, the preset accident characteristics for each type include the age A and gender G of the accident participants (male = 1, female = 0), the weather conditions W (sunny = 1, other = 0), the traffic signal control at the accident location K (signal control present = 1, other = 0), and whether the accident location is an intersection C (yes = 1, no = 0). The traffic signal control at the accident location includes whether there are traffic lights or not. The preset road characteristics for each type include road network density T, road land area D, green land area L, main road density M, and secondary road density B.
[0043] In step A, specifically through the following steps, a latent category classification is used to construct a clustering model, dividing each traffic accident into a preset number of clusters:
[0044] Step A1: For each traffic accident in the study area, based on the preset accident characteristics and road characteristics corresponding to each traffic accident, and combined with the preset number of clusters, construct the initial clustering model as shown below:
[0045]
[0046] In the formula, p(c|x i x represents the probability that traffic accident i belongs to group c, where c∈C; i This represents the preset accident characteristics corresponding to traffic accident i and the preset road characteristics corresponding to the study area; C is the total number of populations; α s α represents the regression parameter corresponding to population s. s This represents the constant corresponding to the population s;
[0047] Step A2: Based on the initial clustering model and a preset probability threshold, classify each traffic accident into a different cluster.
[0048] Furthermore, the number of groups in the preset number of groups is obtained by iteratively performing the following steps:
[0049] Step 1: For each traffic accident in the study area, based on the preset accident characteristics and road characteristics corresponding to each traffic accident, and combined with the number of clusters (the initial number of clusters is 1), construct a clustering model.
[0050] Step 2: Based on the constructed population clustering model, obtain AIC and BIC;
[0051]
[0052]
[0053] In the formula, SSE represents the sum of squared errors for each cluster in the clustering model; I represents the total number of traffic accidents in the study area; L represents the total number of preset accident features and preset road features for each type.
[0054] Step 3: Based on AIC and BIC, and combined with AIC and BIC from the previous iteration, if the percentage change in AIC and BIC is less than the preset percentage threshold, then the number of populations corresponding to the current iteration is the preset number of populations, and the iteration ends; otherwise, increment the number of populations by one and return to Step 1.
[0055] Step B: Based on the traffic accidents corresponding to each ethnic group, and combined with the pre-set road characteristics of each type in the study area, construct a safety analysis model corresponding to each ethnic group; that is, take the pre-set road characteristics of each type as safety influencing factors.
[0056] The security analysis models for each ethnic group in step B are obtained through the following process:
[0057] Step B1: Based on the traffic accidents corresponding to each ethnic group, and combined with the pre-defined road characteristics of each type in the study area, a safety analysis model corresponding to each ethnic group is constructed using a negative binomial regression model with random parameters. The model is shown in the following formula.
[0058]
[0059] In the formula; N ci X represents the total number of traffic accidents in group c; j Let ε represent the feature data of the j-th type of road in the group, ε be the regression model error corresponding to the group, β0 represent the constant corresponding to the group, and β be the feature data of the j-th type of road in the group. j is the regression parameter corresponding to the j-th type of road feature in the population; J represents the preset total number of road features of each type.
[0060] Step B2: For the security analysis model, based on β j The security analysis model has been revised and updated. The updated security analysis model is shown below:
[0061]
[0062] In the formula, It follows a normal distribution with a mean of 1 and a variance of δ. 2 Among them, β j By introducing This can further improve the model's performance in exploring heterogeneity.
[0063] Step C: Based on the safety analysis model corresponding to each ethnic group, obtain the characteristics of each type of road that affect road safety for each ethnic group, that is, obtain the safety influencing factors that affect road safety, thereby realizing the safety evaluation of roads in the study area, that is, constructing a safety interpretation of roads in the study area based on the preset characteristics of each type of road.
[0064] In step C, based on the safety analysis models corresponding to each ethnic group, the characteristics of various types of roads affecting road safety in the target area are obtained through the following process, thereby achieving a safety evaluation of the roads in the study area:
[0065] When β′ j If β' > 0, it indicates that the characteristics of road type j have a positive impact on the occurrence of accidents in the population. j A value less than 0 indicates that the j-th type of road feature has a negative impact on the occurrence of accidents within the group. The safety factors affecting road safety are the road features of each type that have a negative impact.
[0066] The following specific embodiments will further illustrate this solution.
[0067] (1) Multi-source traffic data collection: Multi-source data was collected through accurate survey methods and relevant departmental research. The data collection table is shown in Table 1 below.
[0068] Table 1
[0069] sample L T M D A W G K C B <![CDATA[N1]]> <![CDATA[L1]]> <![CDATA[T1]]> <![CDATA[M1]]> <![CDATA[D1]]> <![CDATA[A1]]> <![CDATA[W1]]> <![CDATA[G1]]> <![CDATA[K1]]> <![CDATA[C1]]> <![CDATA[B1]]> <![CDATA[N2]]> <![CDATA[L2]]> <![CDATA[T2]]> <![CDATA[M2]]> <![CDATA[D2]]> <![CDATA[A2]]> <![CDATA[W2]]> <![CDATA[G2]]> <![CDATA[K2]]> <![CDATA[C2]]> <![CDATA[B2]]> ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ <![CDATA[N 10 ]]> <![CDATA[L 10 ]]> <![CDATA[T 10 ]]> <![CDATA[M 10 ]]> <![CDATA[D 10 ]]> <![CDATA[A 10 ]]> <![CDATA[W 10 ]]> <![CDATA[G 10 ]]> <![CDATA[K 10 ]]> <![CDATA[C 10 ]]> <![CDATA[B 10 ]]> <![CDATA[N 11 ]]> <![CDATA[L 11 ]]> <![CDATA[T 11 ]]> <![CDATA[M 11 ]]> <![CDATA[D 11 ]]> <![CDATA[A 11 ]]> <![CDATA[W 11 ]]> <![CDATA[G 11 ]]> <![CDATA[K 11 ]]> <![CDATA[C 11 ]]> <![CDATA[B 11 ]]> <![CDATA[N 12 ]]> <![CDATA[L 12 ]]> <![CDATA[T 12 ]]> <![CDATA[M 12 ]]> <![CDATA[D 12 ]]> <![CDATA[A 12 ]]> <![CDATA[W 12 ]]> <![CDATA[G 12 ]]> <![CDATA[K 12 ]]> <![CDATA[C 12 ]]> <![CDATA[B 12 ]]> ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ ~ <![CDATA[N 100 ]]> <![CDATA[L i ]]> <![CDATA[T i ]]> <![CDATA[M i ]]> <![CDATA[D i ]]> <![CDATA[A i ]]> <![CDATA[W i ]]> <![CDATA[G i ]]> <![CDATA[K i ]]> <![CDATA[C i ]]> <![CDATA[B i ]]>
[0070] The collected traffic accident sample feature data includes the preset accident characteristics for each type, including the age A and gender G of the accident participants (male = 1, female = 0), the weather conditions W (sunny = 1, other = 0), the traffic signal control at the accident location K (signal control present = 1, other = 0), and whether the accident location is an intersection C (yes = 1, no = 0). The traffic signal control at the accident location is defined as whether there are traffic lights or not. The preset road characteristics for each type include road network density T, road land area D, green land area L, main road density M, and secondary road density B.
[0071] (2) Traffic accidents are divided into different groups, and the accidents are clustered according to the characteristics of the traffic accident samples. The clustering model is shown below:
[0072]
[0073] Where p(c|x) i x represents the probability that traffic accident i belongs to group c, where c∈C; i This represents the preset accident characteristics and preset road characteristics corresponding to traffic accident i. In this embodiment, it is assumed that the continuous change rate of AIC and BIC is less than 0.01 when the number of groups is 4 (compared to the number of groups is 3). Therefore, the final number of groups in this case is set to 4.
[0074] (3) Constructing safety analysis models for each group: Based on the collected pre-defined road characteristics as safety influencing factors, a random negative binomial regression model was used to construct safety analysis models for each of the four groups:
[0075]
[0076] Where, N ci X represents the total number of traffic accidents in group c; j Let β be the feature data of the j-th type of road in the population. j By introducing This can further improve the model's performance in exploring heterogeneity. Update the security analysis models corresponding to the four groups.
[0077] (4) Analysis of influencing factors: Based on the regression coefficient β in the security analysis model of each ethnic group i The positive or negative value can be used to determine the mechanism of action on the occurrence of accidents, that is, to obtain the safety influencing factors affecting road safety, thereby realizing the safety interpretation of roads in the study area.
[0078] Based on the above methods, this scheme also designs a road safety analysis method based on eliminating the heterogeneity of traffic accident samples, including a clustering module, a safety analysis module, and a safety evaluation module. The clustering module constructs a clustering model for each traffic accident in the study area based on the preset accident characteristics of each type of traffic accident and the preset road characteristics of each type of road in the study area, and divides each traffic accident into a preset number of clusters.
[0079] The safety analysis module constructs a safety analysis model for each group based on the traffic accidents corresponding to each group and the pre-defined road characteristics of each type in the study area.
[0080] The safety evaluation module obtains the road characteristics that affect road safety for each type of road based on the safety analysis model corresponding to each ethnic group.
[0081] A terminal for a road safety analysis method to eliminate heterogeneity in traffic accident samples includes a memory and a processor, which are communicatively connected. The memory stores computer instructions, and the processor executes the computer instructions to perform the road safety analysis method for eliminating heterogeneity in traffic accident samples.
[0082] A computer-readable medium for storing software includes instructions executable by one or more computers, which, when executed by the one or more computers, perform a road safety analysis method for eliminating heterogeneity in traffic accident samples.
[0083] This invention designs a road safety analysis method and system to eliminate the heterogeneity of traffic accident samples. First, accidents are clustered according to their occurrence characteristics, forming groups with different attribute features. Then, based on the safety influencing factors of the accidents, traffic safety models are performed on different groups. A safety analysis model is applied to obtain the influencing factors of road safety in the affected area, and a safety evaluation of the area is conducted. Through the technical solution of this invention, a precise, comprehensive, objective, and truthful road safety evaluation method that reflects the impact data is provided, with a wider range of applications. In particular, clustering based on accident characteristics can amplify the heterogeneity among accident samples, better characterizing the impact of the same influencing factors on different accident groups.
[0084] The above are merely preferred embodiments of the present invention, but do not limit the patent scope of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of the present invention specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the patent protection scope of the present invention.
Claims
1. A road safety analysis method for eliminating heterogeneity in traffic accident samples, characterized in that: For the study area, the following steps are performed to conduct a safety assessment of the roads in the study area: Step A: For each traffic accident in the study area, based on the preset accident characteristics of each type of traffic accident and the preset road characteristics of each type of road in the study area, construct a clustering model to divide each traffic accident into a preset number of clusters. Step B: Based on the traffic accidents corresponding to each ethnic group, and combined with the pre-defined road characteristics of each type in the study area, construct a safety analysis model corresponding to each ethnic group. Step C: Based on the safety analysis model corresponding to each ethnic group, obtain the characteristics of various types of roads that affect road safety for each ethnic group, and realize the safety evaluation of roads in the study area; In step A, a clustering model is constructed through the following steps to divide each traffic accident into a predetermined number of clusters: Step A1: For each traffic accident in the study area, based on the preset accident characteristics and road characteristics corresponding to each traffic accident, and combined with the preset number of clusters, construct the initial clustering model as shown below: ; In the formula, Indicates traffic accident Belonging to an ethnic group The probability, ; Indicates traffic accident The corresponding preset accident characteristics for each type and the preset road characteristics for each type of study area; Total number of ethnic groups; This represents the regression parameters corresponding to population s. This represents the constant corresponding to the population s; Step A2: Based on the initial clustering model and a preset probability threshold, classify each traffic accident into a different cluster.
2. The road safety analysis method for eliminating heterogeneity in traffic accident samples according to claim 1, characterized in that: The number of groups of the preset number of groups is obtained by iteratively executing the following steps: Step 1: For each traffic accident in the study area, based on the preset accident characteristics and road characteristics corresponding to each traffic accident, and combined with the number of clusters (the initial number of clusters is 1), construct a clustering model. Step 2: Based on the constructed population clustering model, obtain AIC and BIC; ; ; In the formula, SSE represents the sum of squared errors for each cluster in the clustering model; I represents the total number of traffic accidents in the study area; L represents the total number of preset accident features and preset road features for each type. Step 3: Based on AIC and BIC, and combined with AIC and BIC from the previous iteration, if the percentage change in AIC and BIC is less than the preset percentage threshold, then the number of populations corresponding to the current iteration is the preset number of populations, and the iteration ends. Otherwise, increment the population count by one and return to step 1.
3. The road safety analysis method for eliminating heterogeneity in traffic accident samples according to claim 1, characterized in that: The security analysis models for each ethnic group in step B are obtained through the following process: Step B1: Based on the traffic accidents corresponding to each ethnic group, and combined with the pre-defined road characteristics of each type in the study area, a safety analysis model corresponding to each ethnic group is constructed using a negative binomial regression model with random parameters. The model is shown in the following formula. ; In the formula; This represents the total number of traffic accidents in group c; For the j-th type of road feature data in the population, The regression model error corresponding to the ethnic group; The constant representing the population; It is the regression parameter corresponding to the j-th type of road feature in the population; This indicates the total number of preset road features for each type; Step B2: For the security analysis model, based on The security analysis model has been revised and updated. The updated security analysis model is shown below: ; In the formula, , It follows a normal distribution with a mean of 1 and a variance of . .
4. The road safety analysis method for eliminating heterogeneity in traffic accident samples according to claim 2, characterized in that: In step C, based on the safety analysis models corresponding to each ethnic group, the characteristics of various types of roads affecting road safety in the target area are obtained through the following process, thereby achieving a safety evaluation of the roads in the study area: when A value greater than 0 indicates that the characteristics of road type j have a positive impact on the occurrence of accidents in the population. If the value is less than 0, it indicates that the characteristics of road type j have a negative impact on the occurrence of accidents in the group.
5. The road safety analysis method for eliminating heterogeneity in traffic accident samples according to claim 1, characterized in that: The preset accident characteristics for each type include the age and gender of the accident participants, the weather conditions at the accident site, the signal control at the accident site, and whether the accident site is an intersection.
6. The road safety analysis method for eliminating heterogeneity in traffic accident samples according to claim 1, characterized in that: The preset road characteristics for each type include road network density, road land area, green land area, main road density, and secondary road density.
7. A system for road safety analysis based on the method for eliminating heterogeneity of traffic accident samples according to any one of claims 1-6, characterized in that: It includes a clustering module, a safety analysis module, and a safety evaluation module. The clustering module targets each traffic accident in the study area and constructs a clustering model based on the preset accident characteristics of each type of traffic accident and the preset road characteristics of each type of road in the study area, dividing each traffic accident into a preset number of clusters. The safety analysis module constructs a safety analysis model for each group based on the traffic accidents corresponding to each group and the pre-defined road characteristics of each type in the study area. The safety evaluation module obtains the road characteristics that affect road safety for each type of road based on the safety analysis model corresponding to each ethnic group.
8. A terminal for a road safety analysis method that eliminates heterogeneity in traffic accident samples, characterized in that: The method includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes the computer instructions to perform the road safety analysis method for eliminating heterogeneity of traffic accident samples as described in any one of claims 1-6.
9. A computer-readable medium for storing software, characterized in that: The readable medium includes instructions executable by one or more computers, which, when executed by the one or more computers, perform a road safety analysis method for eliminating heterogeneity in traffic accident samples as described in any one of claims 1-6.