Method, device and product for predicting occurrence probability of traffic accident on urban express highway

The method integrates traffic flow and GPS data with DS evidence theory and machine learning algorithms to optimize traffic accident prediction models, addressing the limitations of existing methods by providing real-time accident probability predictions and warnings.

JP2025179776AActive Publication Date: 2025-12-10TAIZHOU UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024128076
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-28
Filing Date
2024-08-02
Publication Date
2025-12-10
Estimated Expiration
2044-08-02

AI Technical Summary

Technical Problem

Existing methods for predicting traffic accidents on urban expressways fail to provide real-time accident probability predictions and advance warnings, and do not fully reflect the impact of traffic flow conditions across various spatial regions.

Method used

A method utilizing traffic flow information, GPS trajectory data of taxis, and DS evidence theory to calculate fusion average speed, combined with Balanced Bagging and GBDT-XGBoost algorithms to train an optimized traffic accident risk prediction model for urban expressways.

Benefits of technology

Enables real-time prediction and advance warning of traffic accidents by fully reflecting traffic flow conditions, improving traffic safety through timely management interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025179776000001_ABST
    Figure 2025179776000001_ABST
Patent Text Reader

Abstract

To provide a prediction method, a device, and a product for predicting an occurrence probability of a traffic accident on an urban express highway in a technical field of safe driving.SOLUTION: A prediction method includes: acquiring a fusion average speed based on first characteristic data, second characteristic data, third characteristic data, and fourth characteristic data; inputting the fusion average speed and the four characteristic data into a prediction model to acquire an occurrence probability of a traffic accident; processing the occurrence probability of the traffic accident and the four characteristic data to acquire a training set; training and optimizing a traffic accident risk prediction model for an urban express highway using the training set; acquiring an optimized model; and predicting an occurrence probability of the urban express highway traffic accident using the optimized model. The present invention can fully reflect an influence of traffic flow states in different space ranges on an accident risk, and the present invention can predict the possibility of an accident in real time and warn of the accident in advance.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of safe driving, and in particular to a method, device and product for predicting the probability of traffic accidents occurring on urban expressways. [Background technology]

[0002] Most existing methods for predicting the probability of traffic accidents on urban expressways focus on macroscopic prediction of traffic accidents. They build traffic accident risk prediction models based on cross-sectional microwave detector data and historical data, identify traffic accident hotspots, and explain the relationship between traffic flow characteristics, vehicle characteristics, driver characteristics, road environment, and traffic accident probability in order to improve road safety. Essentially, traditional traffic accident prediction methods do not predict whether or not a traffic accident will occur, but rather estimate the long-term average accident frequency for a particular road. Their drawbacks include an inability to fully reflect the impact of traffic flow conditions across a wide range of spatial regions on accident risk, and an inability to provide real-time predictions of accident probability and advance warning of accidents. Summary of the Invention [Problem to be solved by the invention]

[0003] The objective of the present invention is to provide a method, device and product for predicting the probability of traffic accidents on urban expressways, which can fully reflect the impact of traffic flow conditions in various spatial ranges on the risk of accidents, predict the possibility of accidents in real time, and provide advance warning of accidents. [Means for solving the problem]

[0004] To achieve the above objectives, the present invention provides the following solutions: A method for predicting the probability of a traffic accident occurring on an urban expressway, comprising: Collecting traffic flow information of each lane of a target road according to a predetermined collection frequency to obtain a set of information flows; performing a validity test and missing data estimation on each traffic flow information in the set of information flows to obtain first characteristic data, and the first characteristic data includes valid traffic flow information in the set of information flows and traffic flow information obtained by estimating the valid traffic flow information in the set of information flows; Collect GPS trajectory data of all taxis in operation within a target city within a set time period and accident records of all traffic accidents in the target city within a set time period to obtain a GPS trajectory set and an accident record set, and the target city is a city where a target road exists; Determine second characteristic data based on the accident record set and the GPS trajectory set, the second characteristic data being GPS trajectory data of all operating taxis traveling on the urban expressway within each target time period in the GPS trajectory set, and each target time period being a time period corresponding to each accident record in the accident record set; calculating third characteristic data and fourth characteristic data based on the second characteristic data, the third characteristic data including the average speed of each operating taxi in each road section of each viaduct in the second characteristic data, and the fourth characteristic data including the average speed of each road section of each viaduct; Obtaining a fusion average velocity using DS evidence theory based on the third characteristic data and the fourth characteristic data; Inputting the fusion average speed, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data into a real-time urban expressway traffic accident risk prediction model to obtain the occurrence probability of a traffic accident; Using Balanced Bagging to process the probability of traffic accidents, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data to obtain a training set; The method includes training and optimizing an urban highway traffic accident risk prediction model using the training set, obtaining an optimized model, and using the optimized model to predict the probability of an urban highway traffic accident occurring.

[0005] In addition, performing a validity test and missing data estimation on each traffic flow information in the set of information flows to obtain first characteristic data may specifically include: performing a validity test on each traffic flow information in the set of information flows to obtain valid traffic flow information; performing missing data estimation on the valid traffic flow information to obtain estimated traffic flow information; The method includes determining the valid traffic flow information and the estimated traffic flow information as first characteristic data.

[0006] In addition, determining the second characteristic data based on the accident record set and the GPS trajectory set specifically includes: Obtaining a target trajectory set by deleting GPS trajectory data of taxis currently traveling on urban roads under elevated tracks from the GPS trajectory set; This includes determining the GPS trajectory data of all operating taxis in each target time period in the target trajectory set as the second characteristic data.

[0007] Furthermore, calculating the third characteristic data and the fourth characteristic data based on the second characteristic data specifically includes: Calculate the average speed of each operating taxi in each road section of each viaduct of the second characteristic data based on the second characteristic data and the length of each road section of each viaduct; The method includes calculating an average speed in each road section of each viaduct based on the average speed of each operating taxi in each road section of each viaduct in the second characteristic data.

[0008] In addition, using the training set to train and optimize the urban expressway traffic accident risk prediction model and obtain the optimized model, specifically: Using the training set, train the urban expressway traffic accident risk prediction model constructed using the GBDT algorithm to obtain a trained urban expressway traffic accident risk prediction model; This involves optimizing the trained urban highway traffic accident risk prediction model using the XGBoost algorithm to obtain the optimized model.

[0009] In addition, performing missing data estimation on valid traffic flow information and obtaining estimated traffic flow information specifically includes the following steps: The method includes performing missing data estimation on the valid traffic flow information using a random walk method, and obtaining estimated traffic flow information.

[0010] The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to realize a method for predicting the probability of a traffic accident occurring on the urban expressway.

[0011] A computer program product includes a computer program, which, when executed by a processor, realizes the above method for predicting the probability of a traffic accident occurring on an urban expressway. [Effects of the Invention]

[0012] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects. The present invention performs a validity test and missing data estimation on each traffic flow information in the set of information flows to obtain first characteristic data, determines second characteristic data based on the accident record set and the GPS trajectory set, calculates third characteristic data and fourth characteristic data based on the second characteristic data, obtains a fusion average speed based on the third characteristic data and the fourth characteristic data using DS evidence theory, and inputs the fusion average speed, the first characteristic data, the second characteristic data, the third characteristic data and the fourth characteristic data into a real-time urban expressway traffic accident risk prediction model to obtain the occurrence probability of a traffic accident, and then calculates a balanced Bagging is used to process the probability of traffic accidents, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data to obtain a training set, and the training set is used to train and optimize a traffic accident risk prediction model for urban expressways to obtain an optimized model, and the optimized model is used to predict the probability of traffic accidents on urban expressways. The present invention combines the third characteristic data and the fourth characteristic data, and trains and optimizes the model based on the combined data, traffic flow information, and GPS trajectory information, so as to fully reflect the impact of traffic flow conditions in various spatial ranges on the risk of accidents, and can predict the possibility of accidents in real time based on the optimized model, and provide advance warning of accidents.

[0013] In order to more clearly describe the embodiments of the present invention or the technical solutions of the prior art, the drawings that need to be used in the embodiments are briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative efforts. [Brief explanation of the drawings]

[0014] [Figure 1] 1 is a principle diagram of a method for predicting the probability of traffic accidents on urban expressways provided by an embodiment of the present invention; [Figure 2] 4 is a flowchart of a method for validating data provided by an embodiment of the present invention. [Figure 3]2 is a flowchart of a taxi GPS data deletion method provided by an embodiment of the present invention; [Figure 4] 2 is a flowchart of an accident-free group selection method provided by an embodiment of the present invention; [Figure 5] This is a diagram illustrating the principle of the Balanced Bagging algorithm. [Figure 6] FIG. 1 is a principle diagram of DS evidence theory provided by an embodiment of the present invention. [Figure 7] FIG. 1 is a diagram illustrating the internal structure of a computing device. [Figure 8] 1 is a flowchart of a method for predicting the probability of a traffic accident on an urban expressway provided by an embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION

[0015] The technical solutions in the embodiments of the present invention are clearly and completely described below with reference to the drawings in the embodiments of the present invention, but obviously, the described embodiments are only a part of the embodiments of the present invention, and are not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present invention. In order to make the above objects, features and advantages of the present invention more clear and understandable, the present invention will be described in more detail below with reference to the drawings and specific embodiments.

[0016] This invention provides a method for predicting the probability of traffic accidents on urban expressways to fill the gap in the lack of thorough research on real-time traffic accident prediction. As shown in Figure 8, the method for predicting the probability of traffic accidents on urban expressways includes the following steps: Step 101: Collect traffic flow information of each lane of a target road according to a predetermined collection frequency to obtain a set of information flows.

[0017] Step 102: Perform a validity test and missing data estimation for each traffic flow information in the set of information flows to obtain first characteristic data, which includes valid traffic flow information in the set of information flows and traffic flow information obtained by estimating the valid traffic flow information in the set of information flows (missing traffic flow data obtained by estimating the valid traffic flow information).

[0018] Step 103: Collect the GPS trajectory data of all taxis in operation within the target city within the set time period and the accident records of all traffic accidents within the target city within the set time period to obtain a GPS trajectory set and an accident record set. The target city is the city where the target road is located. The GPS trajectory data includes, but is not limited to, the taxi license plate number, date, time, passenger status, longitude and latitude, speed, and direction.

[0019] Step 104: Determine second characteristic data based on the accident record set and the GPS trajectory set, where the second characteristic data is the GPS trajectory data of all taxis in operation traveling on the urban expressway within each target time period in the GPS trajectory set, and each target time period is a time period corresponding to each accident record in the accident record set.

[0020] Step 105: Calculate third characteristic data and fourth characteristic data based on the second characteristic data, where the third characteristic data includes the average speed of each operating taxi in each road section of each viaduct in the second characteristic data, and the fourth characteristic data includes the average speed of each road section of each viaduct.

[0021] Step 106: Obtain a fusion average velocity using DS evidence theory based on the third characteristic data and the fourth characteristic data.

[0022] Step 107: Input the fused average speed, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data into a real-time urban expressway traffic accident risk prediction model to obtain the probability of traffic accident occurrence.

[0023] Step 108: Use Balanced Bagging to process the traffic accident occurrence probability, the first characteristic data, the second characteristic data, the third characteristic data and the fourth characteristic data to obtain a training set.

[0024] Step 109: Use the training set to train and optimize an urban highway traffic accident risk prediction model, obtain an optimized model, and use the optimized model to predict the probability of urban highway traffic accidents.

[0025] In practical application, collecting traffic flow information of each lane of the target road according to a predetermined collection frequency to obtain a set of information flows is specifically as follows: Side-mounted microwave cross-sectional area sensors are used to collect traffic flow information for each lane of a target road every 30 seconds, with each lane corresponding to multiple side-mounted microwave cross-sectional area sensors. Traffic flow information for each lane includes the device numbers of all side-mounted microwave cross-sectional area sensors in the lane, the lane number, the collection date, the flow rate (the speed of all vehicles passing through the lane within 30 seconds), the speed (the average speed of all vehicles passing each microwave vehicle sensor in the lane within 30 seconds), the collection time, and the time occupancy rate (the ratio of the cumulative time that vehicles passed through all microwave vehicle sensor sensors in the lane within 30 seconds to the total time for 30 seconds).

[0026] In practical application, the raw data collected by the cross-sectional microwave measuring instrument may contain partial errors or missing segments, so it is necessary to test its validity before extracting real-time traffic flow features. Specifically, performing validity testing and missing data estimation on each traffic flow information in the set of information flows to obtain first characteristic data includes: The method further includes performing a validity test on each traffic flow information in the set of information flows to obtain valid traffic flow information. The data validity test is performed according to the data validity test steps shown in Figure 2, and the valid traffic flow data after the test is saved.

[0027] For data that pass the validity test, if the collected data itself is missing, it is necessary to estimate the missing traffic flow data to ensure data completeness.

[0028] In practical applications, much of the traffic flow data corresponding to the extracted accident time and space range is discretely missing, so the random walk method is used to estimate the missing data for valid traffic flow information and obtain the estimated traffic flow information. The random walk formula is as follows:

[0029]

number

[0030] In the formula, X' t+1 is the parameter estimate at time t+1 (specifically, in the present invention, it refers to the traffic flow, speed, or time occupancy rate), and X t is the parameter observation at time t, i.e., the collected value.

[0031] In practical application, collecting all traffic accident records in the target city within a set time period and obtaining a set of GPS trajectories and accident records is specifically as follows: The target city's public security department will obtain all accident records for the past year in the target city. Each accident record will include the date, time, location, driver information, accident information, and a brief description of the accident.

[0032] In practical application, determining the second characteristic data based on the accident record set and the GPS trajectory set specifically includes: This involves removing the GPS trajectory data of taxis in motion traveling on urban roads under elevated tracks from the GPS trajectory set to obtain a target trajectory set. Regarding taxi GPS data, since urban expressways are often located on elevated tracks, when spatially aggregating taxi GPS, it is necessary to remove vehicles traveling on urban roads under elevated tracks in order to extract realistic and effective traffic flow information reflecting taxis in motion. The GPS data of taxis in motion is removed according to Figure 3 to determine the taxis in motion traveling on urban expressways during the time periods corresponding to each accident.

[0033] The GPS trajectory data of all taxis in operation during each target time period of the target trajectory set is determined as second characteristic data. The GPS trajectory data of taxis in operation traveling on an urban expressway within a time range corresponding to each accident is the second characteristic data.

[0034] In practical application, the GPS trajectory data of the taxi can only obtain the GPS position of the taxi and the instantaneous speed of the vehicle, so only the traffic flow speed feature is extracted based on the second characteristic data, and the third characteristic data and the fourth characteristic data are calculated based on the second characteristic data, specifically: Calculating the average speed of each operating taxi on each road section of each viaduct in the second characteristic data based on the second characteristic data and the length of each road section of each viaduct is performed using the following formula:

[0035]

number

[0036] Here, n is the total number of GPS trajectory data of taxi p in operation on road section k of a specific viaduct, Vq represents the speed value of the qth GPS trajectory data of taxi p in operation, L is the length of road section k of the specific viaduct, and as third characteristic data, Vp represents the average speed of taxi p in operation on road section k of the specific viaduct.

[0037] Based on the average speed of each operating taxi in each road section of each viaduct in the second characteristic data, the average speed in each road section of each viaduct is calculated, and the calculation formula is as follows:

[0038]

number

[0039] Here, M is the number of vehicles on the road section k of a specific viaduct, and V k is the average speed of road section k of a specific viaduct.

[0040] In practical application, based on the third and fourth characteristic data, using DS evidence theory to obtain the fusion average speed is specifically: DS evidence theory can build an effective fusion model for heterogeneous sensor data, the principle of which is shown in Figure 6. Part (a) of Figure 6 shows the DS data fusion method, that is, all data is unified and fused using the DS data theory fusion method, and finally the fusion is completed and the data is output. Part (b) of Figure 6 shows another DS data fusion method, which fuses two sets of data respectively, then fuses the fused data with another set of data, and continues until all data is fused, and finally outputs the fused data. The effects of the two methods are equivalent.

[0041] The specific steps are as follows: Step 1: Based on the third and fourth characteristic data, construct the following knowledge identification framework of DS evidence theory.

[0042]

number

[0043] V mircrove (t) and V GPS (t) are derived from the fourth characteristic data and the third characteristic data, respectively, and V mircrove(t) represents the average speed of a specific road section based on a cross-sectional microwave measuring instrument during time period t, i.e., the average speed of a taxi operating on the viaduct road section, and V GPS (t) represents the average speed of a specific road section based on taxi GPS trajectory data during time period t, i.e., the average speed of the road section on the viaduct. Since the two can be considered mutually exclusive, the power set consisting of all elements of F(t) is as follows:

[0044]

number

[0045] where X1(t)=V mircrove indicates that the decision is the average speed of the specific road section of the viaduct of the fourth characteristic data, and X2(t) = V GPS indicates that the decision is the average speed of a taxi in operation on a specific road section of the viaduct of the third characteristic data, and X3(t) = V mircrove (t)∩V GPS (t) is an uncertain decision, meaning that it is impossible to distinguish between the decision X1(t) and X2(t).

[0046] Step 2: Under the two types of data sources, determine the basic trust distribution for different decisions as follows: The distribution of the underlying trust function is a quantitative measure of the degree of support for a particular decision and also the evidence for each decision.

[0047]

number

[0048]

number

[0049] where i represents the evidence provided by the ith data source, j is the jth decision, and m i (Xj (t)) is the number of times the evidence provided by the ith data source is X j (t) indicates the degree to which the decision X is supported. j (t) is called the basic distribution value, P j (X j (t)) is the X in the i-th data source j (t) is the basic probability distribution function, m i (φ) represents the basic distribution value of decision φ, and X j (t) are the three types of decisions in step 1.

[0050] P i (X j (t)) is V i (t)~N(u i (t),σ 2 i Assuming that (t) is satisfied, then, u i (t) and σ 2 i (t) respectively represent the mean and variance of the average speed of the road section in time period t in the historical data of the i-th data source. i (t) represents the average speed of the road section in the historical data of the ith data source in the t-th time period, and P i (X1(t)) represents the basic probability distribution function of X1(t) in the i-th data source, and P i (X2(t)) represents the basic probability distribution function of X1(t) in the i-th data source, and P i (X3(t)) represents the basic probability distribution function of X1(t) for the i-th data source.

[0051] P i (X j The specific calculation method for (t)) is as follows:

[0052]

number

[0053] Step 3: Based on the Demspster evidence of the two data sources, the third characteristic data and the fourth characteristic data, the two data sources are synthesized as follows:

[0054]

number

[0055] Here, K represents the contradiction between the evidence provided by the two data sources, m1(B) represents the basic probability distribution function of the third characteristic data, and m2(C) represents the basic probability distribution function of the fourth characteristic data. The closer 1 / K is to 0, the greater the contradiction between the evidence provided by different data sources. A represents the above uncertain decision, B indicates that the above decision is the average speed of a specific road section in the third characteristic data, and C indicates that the above decision is the average speed of a specific road section in the fourth characteristic data.

[0056] Step 4: According to the speed weight calculation formula, determine the weights of the two types of data sources as follows:

[0057] The formula for calculating the weight based on the average speed during time period t of the third characteristic data is as follows:

[0058]

number

[0059] The weight calculation formula based on the average speed during time period t of the fourth characteristic data is as follows:

[0060]

number

[0061] Finally, by substituting the above formula, we can obtain the average speed for time period t after combining the two data sources as follows:

[0062]

number

[0063] Here, m(X1(t)) represents the basic distribution value of the fourth characteristic data source, m(X2(t)) represents the basic distribution value of the third characteristic data source, ωmircrove is the weight of the third characteristic data, and ωGPS is the weight 1-α(t) of the fourth characteristic data.

[0064] The traffic flow conditions before an accident occur are closely related to the magnitude of accident risk, and are therefore called pre-accident conditions. A traffic condition without an accident under certain conditions is defined as a normal traffic flow condition. To study and determine the changes in traffic flow conditions during an accident and under normal traffic flow conditions, a control group (i.e., data from normal traffic conditions) must be established to compare with the data from the accident group. The control group data are identified according to the non-accident group selection steps shown in Figure 4. Because traffic accidents are low-probability events, the experimental data from the control group will be significantly greater than the experimental data from the accident group. Therefore, creating a real-time urban highway traffic accident risk prediction model becomes an imbalanced data classification problem, i.e., the number of samples from the majority class far exceeds the number of samples from the minority class. Until now, the "case-control" method has been primarily used to address the imbalanced data classification problem in creating traffic accident prediction models. Data from a traffic accident condition is selected as the case, and the corresponding data from a non-accident condition is used as the control. Many subsequent studies have used empirical methods to set the ratio of case group to control group data at 1:4. However, choosing a 1:4 ratio based on empirical methods means that a large amount of control group data is not included in the model, which can impair the model's predictive performance. Therefore, the present invention adopts the Balanced Bagging integrated sampling method to process unbalanced sample sets. Balanced Bagging is used to process traffic accident probability, first characteristic data, second characteristic data, third characteristic data, and fourth characteristic data to obtain a training set. The principle of the Balanced Bagging algorithm is shown in Figure 5. The set consisting of the control group data and the experimental group data is directly divided into subsets containing both the control group data and the experimental group data, and the control group data and the experimental group data are resampled in different subsets to classify the traffic accident probability prediction model and generate multiple classifiers. The results of all the classifiers are then weighted and combined to obtain the final training set.

[0065] In practical application, using the training set to train and optimize a traffic accident risk prediction model for urban expressways and obtaining an optimized model specifically includes the following: The training set is used to train the urban expressway traffic accident risk prediction model constructed using the GBDT algorithm, and a trained urban expressway traffic accident risk prediction model is obtained. The XGBoost algorithm is used to optimize the trained urban highway traffic accident risk prediction model, and the optimized model is obtained.

[0066] In practical application, the steps of using the training set to train the urban expressway traffic accident risk prediction model constructed using the GBDT algorithm and obtaining the trained urban expressway traffic accident risk prediction model are specifically as follows:

[0067] Step 1: Constructing the initial accident risk prediction objective function y0':

[0068]

number

[0069] y i represents the probability of a traffic accident occurring corresponding to the i-th sample group in the training set, where one sample group includes the corresponding probability of a traffic accident occurring, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data; c represents the predicted value obtained by inputting the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data corresponding to the i-th data group into the accident risk prediction model; and h it represents the decision tree model for the t-th iteration, and N represents the total number of groups of samples.

[0070] The loss function can be defined as follows:

[0071] JPEG2025179776000015.jpg1491

[0072] where n represents the total number of samples in the training set, and y' t represents the accident risk prediction objective function for the t-th iteration.

[0073] The accident risk prediction model can be defined as follows:

[0074]

number

[0075] where h it is the decision tree model after the tth iteration.

[0076] Step 2: Train the model for t=1 to T iterations as follows: 1) The residual value r of the accident risk classification (i = 1 to N) for each experimental group and control group at the tth iteration it Calculate y' t-1 represents the accident risk prediction objective function for the t-1th iteration.

[0077]

number

[0078] 2) Using the training set, we create a decision tree model h it Build:

[0079]

number

[0080] 3) Minimize the loss function and set the weight β t Calculate the objective function y t Update '. JPEG2025179776000019.jpg970JPEG2025179776000020.jpg1275JPEG2025179776000021.jpg2876

[0081] 4) Update the iteration number t, proceed to the next iteration, return to 1), repeat until t=T and stop, and then obtain a trained urban highway traffic accident risk prediction model.

[0082] Step 3: Output the trained urban expressway traffic accident risk prediction model.

[0083] The traffic accident risk prediction model for urban expressways established by this invention has good application prospects, and enables early warning of accident risks based on traffic flow parameters, allowing traffic managers to timely issue relevant traffic information to take appropriate safety measures.

[0084] This invention realizes a cross-sectional microwave measuring instrument for measuring traffic flow data that falls into the full-sample detection category. However, it can only collect cross-sectional traffic flow conditions and cannot obtain traffic flow information for the entire road section. Taxi GPS trajectory data can effectively overcome this drawback by continuously tracking traffic information for the entire road section. As shown in Figure 1, this invention proposes integrating two types of data sources using the DS evidence theory method to provide a database for traffic accident risk prediction. A balanced bagging resampling technique was used to balance the accident and non-accident sample data, better resolving the imbalance classification problem associated with real-time accident prediction. Finally, a "black box model" for real-time accident risk prediction on urban expressways was interpreted based on feature importance analysis and partial dependency diagrams, and related analyses were conducted.

[0085] According to the present invention, traffic management departments can receive early warning of the risk of traffic accidents, timely plan traffic police management, and more effectively improve the traffic safety level of urban road networks.

[0086] The present invention can effectively predict the rate of traffic accidents on urban roads, enable traffic management departments to accurately provide early warning of traffic accident risks, formulate traffic management schedules in a timely manner, and avoid traffic accidents, thereby effectively improving the traffic safety level of urban road systems.

[0087] Regarding the creation of the model, the present invention is driven by the XGBoost algorithm, an integrated machine learning algorithm. The model created by the XGBoost algorithm can better fit the nonlinear relationship between real-time traffic flow characteristics and accident risk on urban expressways and has better interpretability.

[0088] In one embodiment, a computer device is further provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for predicting the probability of traffic accidents on urban expressways described in the above method embodiments is realized. The computer device may be a database, the internal structure of which is shown in FIG. 7. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, memory, and I / O interface are connected via a system bus, and the communication interface is connected to the system bus via the I / O interface. The processor of the computer device is used to provide computing and control functions. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for executing the operating system and the computer program stored in the non-volatile storage medium. The database of the computer device is used to store pending transactions. The I / O interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal via a network connection. The computer program, when executed by the processor, implements a data processing method.

[0089] In one embodiment, a computer program product is provided that includes a computer program, which, when executed by a processor, realizes the above-described method for predicting the probability of a traffic accident occurring on an urban expressway described in the above-described method embodiment.

[0090] It is necessary to explain that all subject information (including but not limited to subject device information, subject personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) related to this application are information and data approved by the subject or fully approved by all parties, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.

[0091] The technical features of the above embodiments can be combined in any manner, and for the sake of brevity, not all possible combinations of each technical feature in the above embodiments are described, but as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0092] In this specification, specific examples are used to explain the principle and implementation method of the present invention, and the description of the above embodiments is only used to understand the method of the present invention and its core concept, and those skilled in the art will make changes to the specific implementation and application scope based on the concept of the present invention. In summary, the contents of this specification should not be interpreted as limiting the present invention.

Claims

1. Collecting traffic flow information of each lane of a target road according to a predetermined collection frequency to obtain a set of information flows; performing a validity test and missing data estimation for each traffic flow information in the set of information flows to obtain first characteristic data, wherein the first characteristic data includes valid traffic flow information in the set of information flows and traffic flow information obtained by estimating the valid traffic flow information in the set of information flows; The GPS trajectory data of all taxis in operation within a target city within a set time period and the accident records of all traffic accidents within the target city within a set time period are collected to obtain a GPS trajectory set and an accident record set, and the target city is a city where a target road exists; determining second characteristic data based on the accident record set and the GPS trajectory set, the second characteristic data being GPS trajectory data of all taxis in operation traveling on the urban expressway within each target time period in the GPS trajectory set, and each target time period being a time period corresponding to each accident record in the accident record set; calculating third characteristic data and fourth characteristic data based on the second characteristic data, wherein the third characteristic data includes an average speed of each operating taxi in each road section of each viaduct in the second characteristic data, and the fourth characteristic data includes an average speed of each road section of each viaduct; Obtaining a fusion mean velocity using D-S evidence theory based on the third characteristic data and the fourth characteristic data; Inputting the fusion average speed, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data into a real-time urban expressway traffic accident risk prediction model to obtain the occurrence probability of a traffic accident; Using balanced bagging, process the traffic accident probability, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data to obtain a training set; A method for predicting the probability of traffic accidents on urban expressways, comprising: training and optimizing a traffic accident risk prediction model for urban expressways using a training set, obtaining an optimized model, and predicting the probability of traffic accidents on urban expressways using the optimized model.

2. Specifically, performing a validity test and missing data estimation on each traffic flow information in the set of information flows to obtain first characteristic data includes: performing a validity test on each traffic flow information in the set of information flows to obtain valid traffic flow information; performing missing data estimation on the valid traffic flow information to obtain estimated traffic flow information; 2. A method for predicting the probability of a traffic accident occurring on an urban expressway as described in claim 1, characterized in that it includes determining the valid traffic flow information and the estimated traffic flow information as first characteristic data.

3. Specifically, determining the second characteristic data based on the accident record set and the GPS trajectory set includes: Obtaining a target trajectory set by deleting GPS trajectory data of taxis currently traveling on urban roads under elevated tracks from the GPS trajectory set; A method for predicting the probability of traffic accidents occurring on urban expressways as described in claim 1, characterized in that it includes determining the GPS trajectory data of all taxis in operation during each target time period in the target trajectory set as the second characteristic data.

4. Specifically, calculating the third characteristic data and the fourth characteristic data based on the second characteristic data includes: Calculating an average speed of each operating taxi in each road section of each viaduct in the second characteristic data based on the second characteristic data and the length of each road section of each viaduct; 2. A method for predicting the probability of a traffic accident occurring on an urban expressway as described in claim 1, further comprising: calculating the average speed of each road section of each viaduct based on the average speed of each operating taxi in each road section of each viaduct in the second characteristic data.

5. Specifically, using the training set to train and optimize the urban expressway traffic accident risk prediction model and obtain the optimized model includes: Using the training set, train the urban expressway traffic accident risk prediction model constructed using the GBDT algorithm to obtain a trained urban expressway traffic accident risk prediction model; 2. The method for predicting the probability of traffic accidents on urban expressways as described in claim 1, further comprising: optimizing the trained urban expressway traffic accident risk prediction model using the XGBoost algorithm to obtain an optimized model.

6. Specifically, performing missing data estimation on valid traffic flow information and obtaining estimated traffic flow information includes: A method for predicting the probability of traffic accidents occurring on urban expressways as described in claim 2, characterized in that it includes using the random walk method to perform missing data estimation on valid traffic flow information and obtaining the estimated traffic flow information.

7. A computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to realize a method for predicting the probability of a traffic accident occurring on an urban expressway according to any one of claims 1 to 6.

8. A computer program product including a computer program, wherein when the computer program is executed by a processor, the computer program product realizes the method for predicting the probability of a traffic accident occurring on an urban expressway according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Accident prediction method, computer program, accident prediction device, and learning model generation method

    JP2021182189A

  • Traffic control device and learning model production method

    JP2022074223A

  • Information processing device

    JP2023086112A