Method for analyzing severity of traffic accidents based on hierarchical clustering evaluation model
By using hierarchical clustering evaluation models and OLS models, the problem of incomplete research on traffic accident risk levels has been solved, enabling a comprehensive classification of the severity of traffic accidents and analysis of influencing factors, thus supporting the formulation of traffic management strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-24
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies have the drawback of being incomplete in studying traffic accident risk levels and cannot effectively conduct cluster analysis of traffic accident risk levels across regions.
A hierarchical clustering-based evaluation model was adopted. The number of minor injuries, serious injuries, and deaths were selected as evaluation indicators to construct an evaluation indicator matrix. The matrix was then dimensionless, and Euclidean distance was used to calculate similarity. Clusters with the lowest similarity were merged, and the silhouette coefficient was used to determine the final number of clusters. In conjunction with the OLS model, influencing factors were analyzed to classify the severity of traffic accidents.
It enables a comprehensive classification and temporal variation analysis of the severity of daily traffic accidents in a region, providing reference conclusions for traffic management strategies.
Smart Images

Figure QLYQS_1 
Figure QLYQS_4 
Figure QLYQS_10
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of traffic accident severity analysis, and in particular to a traffic accident severity analysis method based on a hierarchical clustering evaluation model. BACKGROUND
[0002] China is a vast country with different traffic environments in different regions. To explore the changes in traffic risk, it is necessary to cluster the traffic accident risk levels. Current research on traffic accident risk levels focuses on the interaction between drivers and the environment, but this research method has the disadvantage of being incomplete. Therefore, we propose a traffic accident severity analysis method based on a hierarchical clustering evaluation model. SUMMARY
[0003] Based on the technical problems existing in the background art, the present application proposes a traffic accident severity analysis method based on a hierarchical clustering evaluation model.
[0004] The traffic accident severity analysis method based on a hierarchical clustering evaluation model proposed by the present application includes the following steps:
[0005] S1: Select the number of light injuries, severe injuries, and deaths in the accident as evaluation indicators, and then construct an evaluation indicator matrix according to the following equation, which is shown below:
[0006] (1),
[0007] In equation (1), m represents the number of evaluation objects, such as representing 2342 days from January 1, 2017 to May 31, 2023, with a value range of n represents the number of evaluation indicators, with a value range of representing the number of light injuries, severe injuries, and deaths, respectively;
[0008] S2: Assign weights to the evaluation indicator matrix in S1, and then use the standardization method to process it to be dimensionless, as shown in the following equation:
[0009] (2),
[0010] represents the mth row and nth column element in the matrix after assigning weights, , represent the maximum element and the minimum element in the matrix after assigning weights, respectively;
[0011] S3: The evaluation indicator matrix after dimensionless processing in S2 is as follows:
[0012] (3),
[0013] Each row of elements in formula (3) is regarded as a cluster, the matrix in formula (1) is regarded as an initial cluster matrix and an initial similarity matrix, the similarity between elements is calculated using the Euclidean distance, and the Euclidean distance formula is as follows:
[0014] (4),
[0015] , respectively represent the value of elements x and y in the n th dimension, wherein n = 3;
[0016] S4: Then the two clusters with the smallest similarity are merged to form a new cluster by the similarity matrix, a new cluster matrix is formed, the similarity of the new cluster matrix is calculated, and the similarity matrix is updated, and the process is repeated until the final clustering number is reached;
[0017] S5: The final clustering number is determined by using the silhouette coefficient, the final clustering number is obtained from 2 to 10 by using a loop function, the corresponding silhouette coefficient is obtained, the maximum silhouette coefficient is selected as the severity classification number, and the calculation formula of the silhouette coefficient is as follows:
[0018] (5),
[0019] Then the silhouette coefficient of all elements is calculated, and the average value is taken, so that the silhouette coefficient of the entire clustering result is obtained, and the traffic accident severity classification is completed;
[0020] S6: The OLS model assumes that there is a linear relationship between the dependent variable and the independent variable, and the linear equation is as follows:
[0021] (6),
[0022] Then the OLS model uses the least square sum of the residuals between the observed values and the predicted values of the model to estimate the optimal regression coefficient of the model, so as to achieve a suitable model fitting effect, and the residual square sum is as follows:
[0023] (7),
[0024] After the OLS model fitting is completed, evaluation analysis is carried out.
[0025] Preferably, in S5, a higher silhouette coefficient can represent a better clustering result.
[0026] Preferably, in S5, the silhouette coefficient of element i in formula (5) is represents the average distance of element i to elements in other clusters. represents the average distance of element i to elements in other clusters.
[0027] Preferably, in the S6, represents the severity of the accident on the first day, is the severity of the accident on the first day, is the regression coefficient, represents the severity of the accident on the first day, is the value of the first influencing factor on the first day, is the value of the first influencing factor on the first day, represents the error term.
[0028] Preferably, in the S6, represents the optimal parameters of the model, represents the sum of squares of residuals, where y is the observed value of the dependent variable, represents the predicted value of the model for the dependent variable.
[0029] Compared with the prior art, the present application aims to classify the severity of traffic accidents in each region every day and analyze the time variation and influencing factors of the severity, so that comprehensive and corresponding severity variation and influencing factor analysis conclusions can be obtained, and the severity variation and influencing factor analysis conclusions can be used as reference materials for traffic management strategies under major public time. DETAILED DESCRIPTION
[0030] The present application will be further described below in conjunction with specific embodiments.
[0031] EMBODIMENT
[0032] The present embodiment proposes a traffic accident severity analysis method based on a hierarchical clustering evaluation model, including the following steps:
[0033] S1: Select the number of light injuries, severe injuries, and deaths in the accident as evaluation indexes, and then construct an evaluation index matrix according to the following equation, which is shown as follows:
[0034] (1),
[0035] In equation (1), m represents the number of evaluation objects, such as representing 2342 days from January 1, 2017 to May 31, 2023, and the value range is , n represents the number of evaluation indexes, and the value range , respectively represent the number of light injuries, severe injuries, and deaths;
[0036] S2: Assign weights to the evaluation index matrix in S1, and then use the standardization method to process it dimensionless, as shown in the following equation:
[0037] (2),
[0038] denotes the mth row and nth column element in the matrix after assigning weights, 、 denotes the maximum element and the minimum element in the matrix after assigning weights, respectively;
[0039] S3: After the dimensionless processing in S2, the evaluation index matrix is as follows:
[0040] (3),
[0041] Each row element in formula (3) is regarded as a cluster, the matrix in formula (1) is regarded as an initial cluster matrix and an initial similarity matrix, the similarity between elements is calculated using the Euclidean distance, and the Euclidean distance formula is as follows:
[0042] (4),
[0043] , denote the values of elements x and y in the nth dimension, respectively, where n=3;
[0044] S4: Then, the two clusters with the smallest similarity are merged to form a new cluster through the similarity matrix, a new cluster matrix is formed, the similarity of the new cluster matrix is calculated, the similarity matrix is updated, and the process is repeated until the final clustering number is reached;
[0045] S5: The final clustering number is determined by using the silhouette coefficient, the final clustering number is hierarchically clustered from 2 to 10 using a loop function, the corresponding silhouette coefficient is obtained, the maximum silhouette coefficient is selected as the severity classification number, a higher silhouette coefficient can indicate a better clustering result, and the calculation formula of the silhouette coefficient is as follows:
[0046] (5),
[0047] In formula (5), is the silhouette coefficient of element i, denotes the distance from element i to other elements in the same cluster, denotes the average distance from element i to elements in other clusters, then by calculating the silhouette coefficients of all elements and taking their average, the silhouette coefficient of the entire clustering result can be obtained, and the traffic accident severity classification is completed;
[0048] S6: The OLS model assumes that there is a linear relationship between the dependent variable and the independent variable, and the linear equation is as follows:
[0049] (6),
[0050] In formula (6) represents the severity of the accident on the first day, is a regression coefficient, represents the value of the first influencing factor on the first day, represents the value of the first influencing factor on the first day, represents an error term, and then the OLS model uses the residual sum of squares between the observed values and the predicted values of the model to estimate the optimal regression coefficient of the model, so as to achieve a suitable model fitting effect, and the residual sum of squares is represented as follows:
[0051] (7),
[0052] In formula (7) represents the optimal parameter of the model, represents the residual sum of squares, wherein y is the observed value of the dependent variable, represents the predicted value of the model for the dependent variable, and evaluation analysis is performed after the OLS model fitting is completed.
[0053] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any person skilled in the art can make equivalent replacement or change according to the technical solution and the inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A method for analyzing the severity of traffic accidents based on a hierarchical clustering evaluation model, characterized in that, Includes the following steps: S1: Select the number of minor injuries, serious injuries, and deaths in the accident as evaluation indicators, and then construct the evaluation indicator matrix according to the following equation. The evaluation indicator matrix is shown below: (1), In equation (1), m represents the number of evaluation objects, and its value range is... n represents the number of evaluation indicators, and its value range is... These represent the number of people with minor injuries, serious injuries, and deaths, respectively. S2: Assign weights to the evaluation index matrix in S1, and then use a standardization method to make it dimensionless, as shown in the following formula: (2), Indicates the weights assigned The element in the m-th row and n-th column of the matrix, , These represent the weights assigned respectively. The maximum and minimum elements in the matrix; S3: The evaluation index matrix after dimensionless processing in S2 is as follows: (3), Each row of elements in equation (3) is considered as a cluster, and the matrix in equation (1) is considered as the initial cluster matrix and the initial similarity matrix. The similarity between elements is calculated using Euclidean distance, and the Euclidean distance formula is as follows: (4), , Let x and y represent the values of elements x and y in the nth dimension, respectively, where n=3; S4: Then, the two clusters with the lowest similarity are merged through the similarity matrix to form a new cluster, forming a new cluster matrix. The similarity of the new cluster matrix is then calculated, and the similarity matrix is updated. This process is repeated until the final number of clusters is reached. S5: The silhouette coefficient is used to determine the final number of clusters. A loop function is used to perform hierarchical clustering from 2 to 10 to obtain the corresponding silhouette coefficients. The cluster with the largest silhouette coefficient is selected as the severity category. The formula for calculating the silhouette coefficient is shown below: (5), Let i be the contour coefficient. This represents the distance from element i to other elements in the same cluster. The silhouette coefficient represents the average distance from element i to elements in other clusters. Then, by calculating the silhouette coefficient of all elements and taking their average value, the silhouette coefficient of the entire clustering result can be obtained. At this point, the classification of the severity of traffic accidents is completed. S6: Using the OLS model, we assume a linear relationship between the dependent and independent variables. The linear equation is shown below: (6) , Representing the The severity of the accident on that day, It is the regression coefficient. Indicates the first The influencing factor in the first The value of the day, Represents the error term. The OLS model then estimates the optimal regression coefficients by minimizing the sum of squared residuals between the observed values and the model predictions, thereby achieving a suitable model fit. The sum of squared residuals is represented as follows: (7), This represents the optimal parameters of the model. Let represent the sum of squared residuals, where y is the observed value of the dependent variable. This represents the model's predicted value for the dependent variable. Evaluation and analysis were performed after the OLS model fitting was completed.
2. The method for analyzing the severity of traffic accidents based on a hierarchical clustering evaluation model according to claim 1, characterized in that, In S5, a higher silhouette coefficient indicates better clustering results.
Citation Information
Patent Citations
Logistics industry development evaluation method and device and electronic system
CN115081981A
Traffic accident severity prediction method and system based on convolutional neural network
CN115689040A