Method for identifying road network accident hotspots based on network kernel density estimation
Patent Information
- Application Number
- CN202610681058.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-18
AI Technical Summary
[0006]为解决上述问题,本申请提供基于网络核密度估计的道路网络事故热点区分识别方法,旨在解决现有技术中无法区分高频热点与高危热点、忽视事故热点时空演变特征和解决核密度值零膨胀和右偏分布导致阈值划分不准确等问题
(1)本发明在传统临界事故率计算架构基础上,引入事故严重程度权重因子,使得不再单一依靠事故频次判别道路热点,同时兼顾事故危害程度,有效规避传统识别方法忽略事故严重度、高危隐患路段漏判的缺陷,提升不良交通风险点位识别的全面性与准确性。
Smart Images

Figure CN122598424A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent transportation engineering, and more specifically, to a method for identifying and distinguishing accident hotspots in road networks based on network kernel density estimation. Background Technology
[0002] Traffic accidents have always been a long-term crisis in the global public safety field, and identifying high-incidence areas is the first step in implementing traffic safety improvement efforts. Therefore, accurately identifying high-incidence areas with key governance value and further distinguishing their types is crucial for developing targeted countermeasures and optimizing resource allocation.
[0003] In recent years, spatial analysis techniques based on Geographic Information Systems (GIS) have provided a powerful analytical framework for hotspot identification, enabling data-driven and targeted intervention in road traffic safety management. Kernel density estimation (KDE) methods, due to their ability to generate continuous risk surfaces and intuitive visualization capabilities, are widely used to detect accident clustering patterns.
[0004] Furthermore, existing technologies use the absolute number of accidents or the cumulative severity as the main indicators for hotspot identification, without taking into account the impact of traffic exposure (such as traffic flow). This leads to high-traffic road sections being mistakenly identified as priority hotspots for governance due to the large absolute number of accidents or the high cumulative severity, equating a large number of accidents with high risk, and ignoring the fact that a large number of accidents may simply be a natural result of a large number of vehicles.
[0005] To achieve the above objectives, this application provides a method for distinguishing and identifying road accident hotspots based on network kernel density estimation. By performing zero-inflation continuous distribution modeling on the network kernel density results after weighting the severity of accidents, and combining spatiotemporal continuity screening and exposure standardization dual-dimensional risk assessment, the method can distinguish and identify high-frequency hotspots and high-risk hotspots, and classify governance priorities. Summary of the Invention
[0006] To address the aforementioned issues, this application provides a method for distinguishing and identifying road network accident hotspots based on network kernel density estimation. This method aims to solve problems in existing technologies, such as the inability to distinguish between high-frequency and high-risk hotspots, the neglect of the spatiotemporal evolution characteristics of accident hotspots, and the inaccurate thresholding caused by zero expansion and right-skewed distribution of kernel density values.
[0007] This invention provides a method for distinguishing and identifying road network accident hotspots based on network kernel density estimation, including: Traffic accident data, road network vector data, and traffic flow data are collected and preprocessed. The severity of the accident is calculated based on the preprocessed data. The data preprocessing includes coordinate calibration, matching of accident points with the road network, and handling of missing values. Based on the unbiased network kernel density estimation method, the severity of the accident is introduced for weighting to obtain a weighted network kernel density estimation model. The optimal bandwidth of the weighted network kernel density estimation model is determined by likelihood cross-validation, and the road network is divided into sub-segments based on the optimal bandwidth. The severity-weighted kernel density value of the sub-road segment is calculated using the weighted kernel density estimation network model. A Hurdle-Gamma model was constructed to fit the severity-weighted kernel density values of each sub-segment and a threshold was used for screening to determine the risk level of each sub-segment. Obtain the risk level of each sub-segment for three consecutive years, and remove sub-segments that have only one year of high risk or are all low risk to obtain a set of sub-segments with spatiotemporal persistence. Spatial aggregation is performed on the spatiotemporal persistent sub-segment set to obtain a candidate hotspot set containing candidate intersection hotspots and candidate road segment hotspots. High-frequency and high-risk hotspots are then identified from the candidate hotspot set.
[0008] In one optional implementation, the severity of the accident is calculated based on the preprocessed data, specifically including: The severity of each accident is calculated using the comprehensive accident intensity method, expressed as follows:
[0009] In the formula, For the accident The severity This represents the total number of deaths in the accident. The number of people seriously injured in the accident. The number of people injured in minor accidents. This refers to the direct property damage caused by the accident.
[0010] In one optional implementation, the weighted network kernel density estimation model includes two diffusion modes: the kernel center is located inside the edge and the kernel center is located on the node. Specifically, it includes: When the kernel center is located inside the edge, the expression for the weighted kernel density estimation model is:
[0011] In the formula, For the density assessment point of the accident The weighted density contribution value, Depending on the severity of the accident, For kernel function, For point Time distance, For density assessment points, The nuclear center is the accident site. For bandwidth, , and From respectively arrive The first, second, and third on the path The degree of each intersection; When the kernel center is located at a node, the expression for the weighted kernel density estimation model is:
[0012] In the formula, The core center is located at the node; The weighted density contribution values of all incidents within the bandwidth are aggregated and normalized to obtain the final severity-weighted density estimate, expressed as follows:
[0013] in,
[0014] In the formula, This is a severity-weighted density estimate. The sum of the severity weights of all observed accidents. This is a collection of all traffic accidents.
[0015] In one optional implementation, the road network is divided into sub-segments based on optimal bandwidth, specifically including: The road network is divided into sub-segments by selecting 1 / 10 of the optimal bandwidth as the dividing criterion. Road segments with a length less than 1 / 10 of the optimal bandwidth are retained as separate sub-segments.
[0016] In one optional implementation, the severity-weighted kernel density value of the sub-road segment is calculated using the weighted kernel density estimation network model, specifically including: The geometric center point of the sub-segment is selected as the density evaluation point of the weighted kernel density estimation model to obtain the severity kernel density value of the sub-segment.
[0017] In one optional implementation, a Hurdle-Gamma model is constructed to fit the severity-weighted kernel density values of each sub-segment and a threshold is used for screening to determine the risk level of each sub-segment, specifically including: The Hurdle-Gamma model includes a logistic regression stage and a Gamma regression stage. In the logistic regression stage, a Logistic model with only an intercept term is used to estimate the probability that the kernel density value is equal to zero or non-zero. In the Gamma regression stage, a Gamma regression model with only an intercept term is used to model the distribution of kernel density values greater than zero. Let the input variable To weight the severity-based kernel density values, a mixture probability density function is constructed based on the two regression stages, expressed as:
[0018] in,
[0019]
[0020] In the formula, For the mixed probability density, In order to observe the global probability of non-zero density, The Gamma density function is a non-zero value. The global mean of positive density values. , The intercept term for the Gamma part. This is the intercept term of the Logistic component. Let Gamma be the shape parameter of the distribution. It is the Gamma function; Integrating the mixed probability density function yields the mixed cumulative distribution function. Under the condition of satisfying the preset cumulative probability, the inverse of the mixed probability density function is obtained to obtain the hotspot threshold. If the mixed probability density value of a sub-segment is greater than the hotspot threshold, it is a high-risk sub-segment; if it is less than or equal to the hotspot threshold, it is a low-risk sub-segment.
[0021] In one optional implementation, the spatiotemporally persistent sub-segment set is spatially aggregated to obtain a candidate hotspot set containing candidate intersection hotspots and candidate road segment hotspots, specifically including: The intersection influence zone is obtained by defining a buffer zone of a predetermined range centered on the intersection node in the road network; Candidate intersection hotspots are those located within the intersection's influence zone, while candidate road segment hotspots are those located outside the intersection's influence zone. These candidate intersection hotspots and candidate road segment hotspots are combined into a candidate hotspot set.
[0022] In one optional implementation, identifying high-frequency and high-risk hotspots from the candidate hotspot set specifically includes: Based on whether they are located in a preset area and whether there is signal control, candidate intersection hotspots are divided into 4 reference groups; based on whether they are located in a preset area, candidate road segment hotspots are divided into 2 reference groups. For candidate intersection hotspots, the exposure of entering vehicles within the statistical period is calculated based on AADT; for candidate road segment hotspots, the exposure of vehicle driving within the statistical period is calculated by combining AADT and road segment length. The actual accident rate is calculated using the following expression:
[0023]
[0024] In the formula, Intersection The actual accident rate, Intersection The number of accidents, Intersection Vehicle exposure levels, For road section The actual accident rate, For road section The number of accidents, For road section Vehicle driving exposure; The critical accident rate is calculated using the following expression:
[0025]
[0026] in,
[0027] In the formula, and Intersections and road sections The critical accident rate, and Intersections and road sections The weighted average accident rate of the reference group. The confidence level coefficient is... and These represent the total number of vehicles entering the intersection and road segment each day; If the actual accident rate of a candidate intersection hotspot or candidate road segment hotspot is greater than the corresponding critical accident rate, then the hotspot is marked as a high-frequency hotspot. The severity-weighted accident rate is calculated using the following expression:
[0028]
[0029] In the formula, and Intersections and road sections The weighted accident rate, and Intersections and road sections The sum of the severity of all accidents; The severity-weighted critical accident rate is calculated using the following expression:
[0030]
[0031] in,
[0032] In the formula, and Intersections and road sections Severity-weighted critical accident rate and Intersections and road sections Severity-weighted average accident rate of the reference group to which it belongs; If the severity-weighted accident rate of a candidate intersection hotspot or candidate road segment hotspot is greater than the corresponding severity-weighted critical accident rate, it is marked as a high-risk hotspot.
[0033] In one optional implementation, the method further includes prioritizing the processing of high-frequency and high-risk hotspots, specifically as follows: If the accident hotspot is both high-frequency and high-risk, it is a composite hotspot and is defined as the first priority. If the accident hotspot is only a high-frequency hotspot or a high-risk hotspot, it is defined as the second priority. Within the same priority level, sort in descending order based on the difference between the actual value and the critical value.
[0034] This application has at least the following advantages or beneficial effects: (1) Based on the traditional critical accident rate calculation framework, this invention introduces an accident severity weight factor, so that the road hotspots are no longer determined solely by the frequency of accidents, but also take into account the degree of accident harm. This effectively avoids the shortcomings of traditional identification methods that ignore the severity of accidents and miss high-risk road sections, and improves the comprehensiveness and accuracy of identifying adverse traffic risk points.
[0035] (2) This invention solves the problems of zero expansion, right skewness, and uneven distribution of road kernel density risk values by constructing a Hurdle Gamma two-stage regression model, and eliminates the interference of skewed data on threshold division.
[0036] (3) This invention uses the dual identification results of the traditional Critical Accident Rate (CCR) and SWCCR methods to classify the identified road hotspots into three categories: high-frequency and high-risk composite hotspots, high-frequency hotspots, and high-risk hotspots. At the same time, a two-level priority ranking mechanism is established to clarify the control priority of different types of hotspots, thus solving the problem that traditional hotspot identification cannot distinguish hotspot types and has ambiguous governance priorities. Attached Figure Description
[0037] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 A flowchart of the road network accident hotspot differentiation and identification method based on network kernel density estimation provided by the present invention; Figure 2 This is a schematic diagram of network kernel density estimation at a node in an embodiment of the present invention. Detailed Implementation
[0039] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0040] Please refer to Figure 1 , Figure 1 This is a flowchart of a road network accident hotspot differentiation and identification method based on network kernel density estimation proposed in an embodiment of this application. Figure 1 As shown, the road network accident hotspot differentiation and identification method based on network kernel density estimation includes: Traffic accident data, road network vector data, and traffic flow data are collected and preprocessed. The severity of the accident is calculated based on the preprocessed data. The data preprocessing includes coordinate calibration, matching of accident points with the road network, and handling of missing values. The calculation of accident severity based on preprocessed data specifically includes: The severity of each accident is calculated using the comprehensive accident intensity method, expressed as follows:
[0041] In the formula, For the accident The severity This represents the total number of deaths in the accident. The number of people seriously injured in the accident. The number of people injured in minor accidents. This refers to the direct property damage caused by the accident.
[0042] In this embodiment, a city in East China is selected for analysis. Traffic accident data is provided by the traffic management department, and data elements include accident number, accident time, location, coordinates, direct property damage, and casualties. Road network vector data is obtained from the OpenStreetMap (OSM) platform and calibrated according to the administrative division layer. The average daily traffic volume of nearly 1,000 intersections and road segments in the city since 2022 is selected. Furthermore, the coordinate system of the map platform was converted from BD09 to WGS84 format to complete coordinate calibration. Since the accident occurred on the road network, a nearest neighbor-based map matching algorithm was used in ArcGIS Pro to project the accident point onto the road network. To address the lack of traffic flow data for unmonitored road sections and intersections, this embodiment conducted supplementary aerial traffic surveys using drone aerial photography to ensure comprehensive data coverage.
[0043] Based on the unbiased network kernel density estimation method, the severity of the accident is introduced for weighting to obtain a weighted network kernel density estimation model. The optimal bandwidth of the weighted network kernel density estimation model is determined by likelihood cross-validation, and the road network is divided into sub-segments based on the optimal bandwidth. Based on the location of the kernel function center in the network, two different diffusion modes are defined: Scenario 1: Kernel center located inside an edge: The kernel function propagates outward along the network path. Whenever the path encounters a node, the kernel value is evenly distributed across all possible unvisited adjacent edges, and so on, until the propagation distance reaches the set bandwidth limit.
[0044] Specifically, the expression for the weighted kernel density estimation model is:
[0045] In the formula, For the density assessment point of the accident The weighted density contribution value, Depending on the severity of the accident, For kernel function, For point Time distance, For density assessment points, The nuclear center is the accident site. For bandwidth, , and From respectively arrive The first, second, and third on the path The degree of each intersection; Scenario 2: The core center is located on a node: such as Figure 2 The diagram shown illustrates the estimation of network kernel density at a node. Main road section , and These are three sub-segments connected to the main path. The kernel function propagates from this node to all connected edges simultaneously, and the kernel value is evenly distributed according to the node's degree (number of connected edges). The kernel value continues to propagate along each path, and the equal distribution operation is performed again when a new node is encountered, until the bandwidth limit is reached.
[0046] Specifically, the expression for the weighted kernel density estimation model is:
[0047] In the formula, The core center is located at the node; Considering two diffusion modes, the weighted density contribution values of all incidents within the bandwidth are aggregated and normalized to obtain the final severity-weighted density estimate, expressed as:
[0048] in,
[0049] In the formula, This is a severity-weighted density estimate. The sum of the severity weights of all observed accidents. This is a collection of all traffic accidents.
[0050] Furthermore, the optimal bandwidth is determined using the likelihood cross-validation method. The higher the log-likelihood value (the closer it is to 0), the better the model fit. In this embodiment, a total of 6 different bandwidths are included, and the likelihood cross-validation for different bandwidths is shown in Table 1: Table 1. Validation results for different bandwidths
[0051] The higher the log-likelihood value (i.e., the closer it is to 0), the better the model fit. In this embodiment, 250m is selected as the optimal bandwidth.
[0052] Furthermore, in order to transform the kernel density function of the continuous network into a comparable discrete spatial analysis unit, the road network is divided into sub-segments. To improve the efficiency of the algorithm, the road network is divided into sub-segments according to 1 / 10 of the optimal bandwidth. For road segments that are less than this length, they are retained as separate sub-units.
[0053] The severity-weighted kernel density value of the sub-road segment is calculated using the weighted kernel density estimation network model. Among them, the geometric center point of each sub-segment is selected as the density evaluation point to obtain the severity-weighted kernel density value of each sub-segment on the road network. This density value is the result of the combined effect of the number and severity of accidents.
[0054] A Hurdle-Gamma model was constructed to fit the severity-weighted kernel density values of each sub-segment and a threshold was used for screening to determine the risk level of each sub-segment. Since the kernel density of a large number of spatial segments in the road network is zero or close to zero, the overall data exhibits a significant right-skewed distribution, which is addressed by constructing a Hurdle-Gamma model.
[0055] The Hurdle-Gamma model consists of two stages, specifically: The first stage, logistic regression: using a logistic model containing only the intercept term to estimate the probability that the kernel density value is equal to zero or non-zero; The second stage, Gamma regression, uses a Gamma regression model containing only the intercept term to model the distribution of kernel density values greater than zero. Gamma regression is naturally applicable to right-skewed distributions and does not require the dependent variable to be an integer, making it more suitable for fitting non-negative continuous data such as kernel density estimates.
[0056] Based on the above two stages, a mixture probability density function is constructed, expressed as follows:
[0057] in,
[0058]
[0059] In the formula, For the mixed probability density, For input variables, In order to observe the global probability of non-zero density, The Gamma density function is a non-zero value. The global mean of positive density values. , The intercept term for the Gamma part. This is the intercept term of the Logistic component. Let Gamma be the shape parameter of the distribution. It is the Gamma function; To derive a classifiable threshold from the output of the Hurdle-Gamma model, the mixture probability density function is integrated to obtain the mixture cumulative distribution function (CDF). Then, under a preset cumulative probability condition, the inverse of the CDF is used to determine the hotspot threshold.
[0060] If the cumulative probability P > the hotspot threshold, it is a high-risk (H) sub-segment; if the cumulative probability P ≤ the hotspot threshold, it is a low-risk (L) sub-segment.
[0061] As alternative implementations, the zero-inflated Poisson model, the zero-inflated negative binomial model, and the Hurdle model can be selected as alternatives to the Hurdle-Gamma model. However, each model is suitable for different situations. The zero-inflated Poisson model is suitable for discrete count data, while the zero-inflated negative binomial model is suitable for excessively discrete data with variance greater than the mean.
[0062] Obtain the risk level of each sub-segment for three consecutive years, and remove sub-segments that have only one year of high risk or are all low risk to obtain a set of sub-segments with spatiotemporal persistence. In this embodiment, the risk level change trend of the same sub-segment over three consecutive years is obtained, and five evolution types are defined: persistent (HHH), emerging (LHH), declining (HHL), fluctuating (HLH), and temporary (H only appears in a certain year). Sub-segment spatiotemporal persistence screening: This invention screens high-risk sub-segments with a certain spatiotemporal persistence, including: persistent, emerging, declining, and fluctuating types. Temporary high-risk sub-segments are excluded at this stage due to their lack of spatiotemporal persistence.
[0063] It should be understood that the screening steps described above are essentially identical in terms of the final screening results; the only difference is in the way they are expressed, and there is no difference in the technical solutions. Among them, the screening based on the evolution type is a concrete classification, which intuitively shows the time-series changes of different risks by enumerating five types of risk evolution patterns: persistent, emerging, declining, volatile, and temporary. Direct screening is a high-level generalization of the screening based on the evolution type.
[0064] Furthermore, in an alternative embodiment, when long-term strategic planning is required and sufficient data is available, data over a longer time span can be used for analysis to classify evolution types, and the specific evolution patterns need to be redefined based on the actual situation. For regions with particularly significant climate change, seasonal evolution analysis can also be used.
[0065] Spatial aggregation is performed on the spatiotemporal persistent sub-road segment set to obtain a candidate hotspot set containing candidate intersection hotspots and candidate road segment hotspots. High-frequency and high-risk hotspots are identified from the candidate hotspot set. Specifically, spatial aggregation is performed on the spatiotemporally persistent sub-road segment set to obtain a candidate hotspot set containing candidate intersection hotspots and candidate road segment hotspots, including: The intersection influence zone is obtained by defining a buffer zone of a predetermined range centered on the intersection node in the road network; Candidate intersection hotspots are located within the intersection's influence zone, while candidate road segment hotspots are located outside the intersection's influence zone.
[0066] In this embodiment, a 250-meter buffer zone centered on the geometric center of the intersection is selected, and the length of the candidate road segment hotspot is preferably 500 meters.
[0067] Among them, candidate intersection hotspots are divided into 4 reference groups based on whether they are located in a preset area and whether there is signal control; candidate road segment hotspots are divided into 2 reference groups based on whether they are located in a preset area. In this embodiment, the reference group classification is shown in Table 2, where INT represents an intersection, SEG represents a road segment, URB represents urban area, SURB represents out-of-town area, SG represents signal control, and ST represents no signal control. Table 2 Reference Group Classification
[0068] Specifically, for candidate intersection hotspots, the exposure of entering vehicles within the statistical period is calculated based on AADT; for candidate road segment hotspots, the exposure of vehicles during the statistical period is calculated by combining AADT and road segment length. The actual accident rate is calculated using the following expression:
[0069]
[0070] In the formula, Intersection The actual accident rate, Intersection The number of accidents, Intersection Vehicle exposure levels, For road section The actual accident rate, For road section The number of accidents, For road section Vehicle driving exposure; The critical accident rate is calculated using the following expression:
[0071]
[0072]
[0073] In the formula, and Intersections and road sections The weighted average accident rate of the reference group. and These represent the total number of vehicles entering the intersection and road segment each day. and Intersections and road sections The critical accident rate, The confidence level coefficient; In this embodiment, The value is 1.645; If the actual accident rate of a candidate hotspot cell is greater than the corresponding critical accident rate, then the candidate hotspot cell is marked as a high-frequency hotspot.
[0074] Identifying high-risk hotspots specifically includes: The severity-weighted accident rate is calculated using the following expression:
[0075]
[0076] In the formula, and Intersections and road sections The weighted accident rate, and Intersections and road sections The sum of the severity of all accidents; The severity-weighted critical accident rate is calculated using the following expression:
[0077]
[0078]
[0079] In the formula, and Intersections and road sections The severity-weighted average accident rate of the reference group to which the accident occurred. and Intersections and road sections Severity-weighted critical accident rate; If the weighted accident rate of a candidate hotspot cell is greater than the corresponding weighted critical accident rate, then the candidate hotspot cell is marked as a high-risk hotspot.
[0080] It needs to be explained that the traditional CCR method only focuses on the number of accidents, implicitly assuming that all accidents cause the same social cost, which does not conform to reality. For example, from the perspective of traffic managers, an area with 3 fatal accidents and an area with 30 accidents that only cause property damage, even if both exceed their respective CCR thresholds, have fundamentally different risk levels.
[0081] Furthermore, the processing priorities for high-frequency and high-risk hotspots are ranked as follows: If the accident hotspot is both high-frequency and high-risk, it is a composite hotspot and is defined as the first priority. If the accident hotspot is only a high-frequency hotspot or a high-risk hotspot, it is defined as the second priority. Within the same priority level, sort in descending order based on the difference between the actual value and the critical value.
[0082] Overall, the road network of this embodiment was analyzed and verified using the method proposed in this invention. The experimental results show that while maintaining a high coverage rate of 80.6%, the spatial cost was reduced to 6.33%, and the PAI value reached 12.74, which is about 99% higher than the cumulative aggregation method (PAI=6.40). Compared with the potential hotspot identification method (PAI=22.72), although the PAI is slightly lower, the coverage rate is increased by 54.44%. Compared with other methods, it achieves a better balance between coverage and spatial cost.
[0083] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0084] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a computer terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0085] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for distinguishing and identifying accident hotspots in road networks based on network kernel density estimation, characterized in that, Including the following steps: Traffic accident data, road network vector data, and traffic flow data are collected and preprocessed. The severity of the accident is calculated based on the preprocessed data. The data preprocessing includes coordinate calibration, matching of accident points with the road network, and handling of missing values. Based on the unbiased network kernel density estimation method, the severity of the accident is introduced for weighting to obtain a weighted network kernel density estimation model. The optimal bandwidth of the weighted network kernel density estimation model is determined by likelihood cross-validation, and the road network is divided into sub-segments based on the optimal bandwidth. The severity-weighted kernel density value of the sub-road segment is calculated using the weighted kernel density estimation network model. A Hurdle-Gamma model was constructed to fit the severity-weighted kernel density values of each sub-segment and a threshold was used for screening to determine the risk level of each sub-segment. Obtain the risk level of each sub-segment for three consecutive years, and remove sub-segments that have only one year of high risk or are all low risk to obtain a set of sub-segments with spatiotemporal persistence. Spatial aggregation is performed on the spatiotemporal persistent sub-segment set to obtain a candidate hotspot set containing candidate intersection hotspots and candidate road segment hotspots. High-frequency and high-risk hotspots are then identified from the candidate hotspot set.
2. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, The severity of the accident is calculated based on the preprocessed data, specifically including: The severity of each accident is calculated using the comprehensive accident intensity method, expressed as follows: In the formula, For the accident The severity This represents the total number of deaths in the accident. The number of people seriously injured in the accident. The number of people injured in minor accidents. This refers to the direct property damage caused by the accident.
3. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, The weighted network kernel density estimation model includes two diffusion modes: the kernel center is located inside the edge and the kernel center is located on the node. Specifically, it includes: When the kernel center is located inside the edge, the expression for the weighted kernel density estimation model is: In the formula, For the density assessment point of the accident The weighted density contribution value, Depending on the severity of the accident, For kernel function, For point Time distance, For density assessment points, The nuclear center is the accident site. For bandwidth, , and From respectively arrive The first, second, and third on the path The degree of each intersection; When the kernel center is located at a node, the expression for the weighted kernel density estimation model is: In the formula, The core center is located at the node; The weighted density contribution values of all incidents within the bandwidth are aggregated and normalized to obtain the final severity-weighted density estimate, expressed as follows: in, In the formula, This is a severity-weighted density estimate. The sum of the severity weights of all observed accidents. This is a collection of all traffic accidents.
4. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, The road network is divided into sub-segments based on optimal bandwidth, specifically including: The road network is divided into sub-segments by selecting 1 / 10 of the optimal bandwidth as the dividing criterion. Road segments with a length less than 1 / 10 of the optimal bandwidth are retained as separate sub-segments.
5. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, The severity-weighted kernel density value of the sub-road segment is calculated using the weighted kernel density estimation network model, specifically including: The geometric center point of the sub-segment is selected as the density evaluation point of the weighted kernel density estimation model to obtain the severity kernel density value of the sub-segment.
6. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, A Hurdle-Gamma model was constructed to fit the severity-weighted kernel density values of each sub-road segment and a threshold was applied to determine the risk level of each sub-road segment. Specifically, this included: The Hurdle-Gamma model includes a logistic regression stage and a Gamma regression stage. In the logistic regression stage, a Logistic model with only an intercept term is used to estimate the probability that the kernel density value is equal to zero or non-zero. In the Gamma regression stage, a Gamma regression model with only an intercept term is used to model the distribution of kernel density values greater than zero. Let the input variable To weight the severity-based kernel density values, a mixture probability density function is constructed based on the two regression stages, expressed as: in, In the formula, For the mixed probability density, In order to observe the global probability of non-zero density, The Gamma density function is a non-zero value. The global mean of positive density values. , The intercept term for the Gamma part. This is the intercept term of the Logistic component. Let Gamma be the shape parameter of the distribution. It is the Gamma function; Integrating the mixed probability density function yields the mixed cumulative distribution function. Under the condition of satisfying the preset cumulative probability, the inverse of the mixed probability density function is obtained to obtain the hotspot threshold. If the mixed probability density value of a sub-segment is greater than the hotspot threshold, it is a high-risk sub-segment; if it is less than or equal to the hotspot threshold, it is a low-risk sub-segment.
7. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, Spatial aggregation is performed on the spatiotemporally persistent sub-segment set to obtain a candidate hotspot set containing candidate intersection hotspots and candidate road segment hotspots, specifically including: The intersection influence zone is obtained by defining a buffer zone of a predetermined range centered on the intersection node in the road network; Candidate intersection hotspots are those located within the intersection's influence zone, while candidate road segment hotspots are those located outside the intersection's influence zone. These candidate intersection hotspots and candidate road segment hotspots are combined into a candidate hotspot set.
8. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, Identifying high-frequency and high-risk hotspots from the candidate hotspot set specifically includes: Based on whether they are located in a preset area and whether there is signal control, candidate intersection hotspots are divided into 4 reference groups; based on whether they are located in a preset area, candidate road segment hotspots are divided into 2 reference groups. For candidate intersection hotspots, the exposure of entering vehicles within the statistical period is calculated based on AADT; for candidate road segment hotspots, the exposure of vehicle driving within the statistical period is calculated by combining AADT and road segment length. The actual accident rate is calculated using the following expression: In the formula, Intersection The actual accident rate, Intersection The number of accidents, Intersection Vehicle exposure levels, For road section The actual accident rate, For road section The number of accidents, For road section Vehicle driving exposure; The critical accident rate is calculated using the following expression: in, In the formula, and Intersections and road sections The critical accident rate, and Intersections and road sections The weighted average accident rate of the reference group. The confidence level coefficient is... and These represent the total number of vehicles entering the intersection and road segment each day; If the actual accident rate of a candidate intersection hotspot or candidate road segment hotspot is greater than the corresponding critical accident rate, then the hotspot is marked as a high-frequency hotspot. The severity-weighted accident rate is calculated using the following expression: In the formula, and Intersections and road sections The weighted accident rate, and Intersections and road sections The sum of the severity of all accidents; The severity-weighted critical accident rate is calculated using the following expression: in, In the formula, and Intersections and road sections Severity-weighted critical accident rate and Intersections and road sections Severity-weighted average accident rate of the reference group to which it belongs; If the severity-weighted accident rate of a candidate intersection hotspot or candidate road segment hotspot is greater than the corresponding severity-weighted critical accident rate, it is marked as a high-risk hotspot.
9. The method for distinguishing and identifying road network accident hotspots based on network kernel density estimation according to claim 1, characterized in that, This also includes prioritizing the processing of high-frequency and high-risk hotspots, specifically: If the accident hotspot is both high-frequency and high-risk, it is a composite hotspot and is defined as the first priority. If the accident hotspot is only a high-frequency hotspot or a high-risk hotspot, it is defined as the second priority. Within the same priority level, sort in descending order based on the difference between the actual value and the critical value.