Method, device, and product for predicting the occurrence probability of traffic accidents on urban expressways

By integrating traffic flow and GPS data with machine learning, the method addresses the limitations of existing accident prediction systems, offering real-time risk assessment and early warnings on urban expressways.

JP7714263B1Active Publication Date: 2025-07-29TAIZHOU UNIV

Patent Information

Application Number
JP2024128076
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2024-05-28
Filing Date
2024-08-02
Publication Date
2025-07-29
Estimated Expiration
2044-08-02

AI Technical Summary

Technical Problem

Existing methods for predicting traffic accidents on urban expressways fail to accurately reflect the influence of traffic flow states in various spatial ranges and cannot provide real-time predictions or warnings before accidents occur.

Method used

A method involving data collection from lane traffic flow, GPS taxi trajectories, and accident records, using D-S evidence theory for data fusion, Balanced Bagging, and machine learning algorithms like GBDT and XGBoost to train a real-time traffic accident risk prediction model.

Benefits of technology

The method provides real-time prediction of accident probabilities and enables early warnings by fully reflecting traffic flow states, enhancing traffic safety through optimized models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714263000001_ABST
    Figure 0007714263000001_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and product for predicting the occurrence probability of traffic accidents on urban expressways, and relates to the technical field of safe driving. 【Solution means】The prediction method includes obtaining a fusion average speed based on the first characteristic data, the second characteristic data, the third characteristic data and the fourth characteristic data, inputting the fusion average speed and the four characteristic data into a prediction model to obtain the occurrence probability of a traffic accident, processing the occurrence probability of the traffic accident and the four characteristic data to obtain a training set, using the training set to train and optimize a traffic accident risk prediction model for urban expressways, obtaining an optimized model, and using the optimized model to predict the occurrence probability of urban expressway traffic accidents. The present invention can fully reflect the influence of traffic flow states in different spatial ranges on accident risks, predict the possibility of accidents in real time, and give early warnings of accidents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of safe driving, and particularly to a method, device, and product for predicting the occurrence probability of traffic accidents on urban expressways.

Background Art

[0002] Many existing methods for predicting the occurrence probability of traffic accidents on urban expressways are macroscopic prediction studies of traffic accidents. Based on data from cross-section microwave detectors and past data, a traffic accident risk prediction model is constructed to identify dangerous locations of traffic accidents and to explain the relationship between characteristics of traffic flow, vehicle characteristics, driver characteristics, road environment, etc. and the occurrence probability of traffic accidents in order to improve the safety of road traffic. Essentially, conventional traffic accident prediction is not to predict the occurrence or non-occurrence of traffic accidents, but to estimate the long-term average accident frequency on a specific road. Its disadvantages are that it cannot fully reflect the influence of traffic flow states in various spatial ranges on accident risks, and at the same time, it cannot predict the possibility of accidents in real time and give a warning before the occurrence of accidents.

Summary of the Invention

Problems to be Solved by the Invention

[0003] The object of the present invention is to provide a method, device, and product for predicting the occurrence probability of traffic accidents on urban expressways, which can fully reflect the influence of traffic flow states in various spatial ranges on accident risks, predict the possibility of accidents in real time, and give a warning before the occurrence of accidents.

Means for Solving the Problems

[0004] To achieve the above object, the present invention provides the following solutions. A method for predicting the occurrence probability of traffic accidents on urban expressways, comprising: collecting traffic flow information of each lane of a target road according to a predetermined collection frequency to obtain a set of information flows; Perform validity tests and missing data estimation on each traffic flow information within the set of the information flows to obtain first characteristic data, where the first characteristic data includes valid traffic flow information within the set of the information flows and traffic flow information obtained by estimating valid traffic flow information within the set of the information flows, and Collect GPS trajectory data of all in-operation taxis within the target city during the set time period and accident records of all traffic accidents in the target city during the set time period to obtain a GPS trajectory set and an accident record set, where the target city is a city where the target road exists, and Determine second characteristic data based on the accident record set and the GPS trajectory set, where the second characteristic data is GPS trajectory data of all in-operation taxis traveling on the urban highway within each target time period in the GPS trajectory set, and each target time period is the time period corresponding to each accident record in the accident record set, and Calculate third characteristic data and fourth characteristic data based on the second characteristic data, where the third characteristic data includes the average speed of each in-operation taxi in each road section of each viaduct within the second characteristic data, and the fourth characteristic data includes the average speed of each road section of each viaduct, and Obtain a fused average speed using the D-S evidence theory based on the third characteristic data and the fourth characteristic data, and Input the fused average speed, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data into a real-time urban highway traffic accident risk prediction model to obtain the occurrence probability of a traffic accident, and Use Balanced Bagging to process the occurrence probability of a traffic accident, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data to obtain a training set, and Train and optimize an urban highway traffic accident risk prediction model using the training set to obtain an optimized model, and use the optimized model to predict the occurrence probability of an urban highway traffic accident.

[0005] Moreover, performing validity tests and missing data estimation on each traffic flow information within the set of the information flows to obtain first characteristic data specifically includes Performing a validity test on each traffic flow information within the set of the information flows to obtain traffic flow information having validity; Performing missing data estimation on the traffic flow information having validity to obtain the traffic flow information obtained by estimation; Determining the traffic flow information having validity and the traffic flow information obtained by estimation as first characteristic data is included.

[0006] In addition, determining second characteristic data based on the accident record set and the GPS trajectory set specifically includes: Deleting the GPS trajectory data of the operating taxis traveling on the urban roads under the viaduct in the GPS trajectory set to obtain a target trajectory set; Determining the GPS trajectory data of all operating taxis in each target time period in the target trajectory set as second characteristic data is included.

[0007] In addition, calculating third characteristic data and fourth characteristic data based on the second characteristic data specifically includes: Calculating the average speed of each operating taxi in each road section of each viaduct of the second characteristic data based on the second characteristic data and the length of each road section of each viaduct; Calculating the average speed in each road section of each viaduct based on the average speed of each operating taxi in each road section of each viaduct of the second characteristic data is included.

[0008] In addition, training and optimizing a traffic accident risk prediction model for urban expressways using a training set to obtain an optimized model specifically includes: Training a traffic accident risk prediction model for urban expressways constructed using the GBDT algorithm using the training set to obtain a trained traffic accident risk prediction model for urban expressways; Optimizing the traffic accident risk prediction model for urban expressways trained using the XGBoost algorithm to obtain an optimized model is included.

[0009] In addition, performing missing data estimation on traffic flow information having validity and obtaining the traffic flow information obtained by estimation specifically includes performing missing data estimation on traffic flow information having validity using a random walk method and obtaining the traffic flow information obtained by estimation.

[0010] A computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method for predicting the occurrence probability of traffic accidents on the urban expressway.

[0011] A computer program product including a computer program, and when the computer program is executed by a processor, implementing the above prediction method for the occurrence probability of traffic accidents on the urban expressway.

Advantages of the Invention

[0012] According to a specific embodiment provided by the present invention, the present invention discloses the following technical effects. The present invention performs validity tests and missing data estimations on each traffic flow information within the set of the information flows to obtain first characteristic data, determines second characteristic data based on the accident record set and the GPS trajectory set, calculates third characteristic data and fourth characteristic data based on the second characteristic data, obtains a fusion average speed using the D-S evidence theory based on the third characteristic data and the fourth characteristic data, inputs the fusion average speed, the first characteristic data, the second characteristic data, the third characteristic data and the fourth characteristic data into a real-time urban highway traffic accident risk prediction model to obtain the occurrence probability of a traffic accident, uses Balanced Bagging to process the occurrence probability of a traffic accident, the first characteristic data, the second characteristic data, the third characteristic data and the fourth characteristic data to obtain a training set, uses the training set to train and optimize an urban highway traffic accident risk prediction model to obtain an optimized model, uses the optimized model to predict the occurrence probability of an urban highway traffic accident. The present invention fuses the third characteristic data and the fourth characteristic data, and trains and optimizes the model based on the fused data, traffic flow information and GPS trajectory information, so as to fully reflect the influence of the traffic flow state in various spatial ranges on the accident risk, predict in real time the possibility of an accident occurring based on the optimized model, and can give a warning before the accident.

[0013] To more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings that need to be used in the embodiments are briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without creative efforts.

Brief Description of the Drawings

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Embodiments for Carrying Out the Invention

[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts are included in the protection scope of the present invention. To make the above objects, features, and advantages of the present invention clearer and easier to understand, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0016] The present invention provides a method for predicting the occurrence probability of traffic accidents on urban expressways to make up for the lack of in-depth research on real-time traffic accident prediction. As shown in FIG. 8, the method for predicting the occurrence probability of traffic accidents on urban expressways includes the following steps. Step 101: Collect traffic flow information of each lane of the target road according to a predetermined collection frequency to obtain a set of information flows.

[0017] Step 102: Perform a validity test and missing data estimation on each traffic flow information within the set of information flows to obtain first characteristic data. The first characteristic data includes valid traffic flow information within the set of information flows and traffic flow information obtained by estimating valid traffic flow information within the set of information flows (missing traffic flow data estimated for valid traffic flow information).

[0018] Step 103: Collect the GPS trajectory data of all operating taxis within the target city during the set time period and the accident records of all traffic accidents in the target city during the set time period to obtain a GPS trajectory set and an accident record set. The target city is a city where the target road exists. The GPS trajectory data includes, but is not limited to, the taxi license plate number, date, time, passenger status, longitude and latitude, speed and direction.

[0019] Step 104: Determine second characteristic data based on the accident record set and the GPS trajectory set. The second characteristic data is the GPS trajectory data of all operating taxis traveling on the urban expressway within each target time period in the GPS trajectory set, and each target time period is the time period corresponding to each accident record in the accident record set.

[0020] Step 105: Calculate third characteristic data and fourth characteristic data based on the second characteristic data. The third characteristic data includes the average speed of each operating taxi in each road section of each viaduct within the second characteristic data, and the fourth characteristic data includes the average speed of each road section of each viaduct.

[0021] Step 106: Based on the third characteristic data and the fourth characteristic data, use the D-S evidence theory to obtain a fused average speed.

[0022] Step 107: Input the fused average speed, the first characteristic data, the second characteristic data, the third characteristic data and the fourth characteristic data into a real-time urban expressway traffic accident risk prediction model to obtain the occurrence probability of a traffic accident.

[0023] Step 108: Use Balanced Bagging to process the probability of traffic accident occurrence, first characteristic data, second characteristic data, third characteristic data, and fourth characteristic data to obtain a training set.

[0024] Step 109: Use the training set to train and optimize the traffic accident risk prediction model for urban expressways, obtain the optimized model, and use the optimized model to predict the probability of traffic accident occurrence on urban expressways.

[0025] In actual application, collecting the traffic flow information of each lane of the target road according to a predetermined collection frequency to obtain a set of information flows specifically Using a side-mounted cross-section microwave measuring device, collect the traffic flow information of each lane road surface of the target road every 30 seconds. One lane corresponds to multiple side-mounted cross-section microwave measuring devices. The traffic flow information of one lane road surface includes the device numbers, lane numbers, collection dates, traffic flow (the speeds of all vehicles passing through the lane within 30 seconds), speeds (the average value of the speeds of all vehicles passing through each microwave vehicle measuring device in the lane within 30 seconds), collection times, and time occupancy rates (the ratio of the cumulative time when vehicles pass through all microwave vehicle measuring device cross-sections in the lane within 30 seconds to 30 seconds) of all side-mounted cross-section microwave measuring devices of the lane.

[0026] In actual application, since the raw data collected by the cross-section microwave measuring device may contain partial errors or missing segments, it is necessary to test its effectiveness before extracting real-time traffic flow characteristics. Conducting an effectiveness test and estimating missing data for each traffic flow information within the set of information flows to obtain first characteristic data specifically Includes obtaining traffic flow information with effectiveness by conducting an effectiveness test on each traffic flow information within the set of information flows. According to the steps of the data effectiveness test shown in Figure 2, the data effectiveness test was executed, and the effective traffic flow data after the test was saved.

[0027] For data that has passed the validity test, if the collected data itself is missing, it is necessary to estimate the missing traffic flow data to ensure data integrity.

[0028] In actual applications, since much of the traffic flow data corresponding to the extracted accident time and space ranges is discretely missing, the random walk method is used to estimate the missing data for valid traffic flow information and obtain the estimated traffic flow information. The formula for the random walk is as follows.

[0029]

Equation

[0030] In the equation, X’ t+1 is the parameter estimated value at time t+1 (in the present invention, specifically referring to traffic flow, speed, or time occupancy), and X t is the parameter observed value at time t, that is, the collected value.

[0031] In actual applications, collecting all accident records of all traffic accidents in the target city within the set time period and obtaining the GPS trajectory set and accident record set specifically means the public security department of the target city obtains all accident records in the target city for the past year. Each accident record includes the date, time, location of occurrence, driver information, accident information, and a brief description of the accident.

[0032] In actual applications, determining the second characteristic data based on the accident record set and GPS trajectory set specifically means It includes deleting the GPS trajectory data of in-service taxis running on urban roads under viaducts in the GPS trajectory set to obtain the target trajectory set. Regarding the GPS data of taxis, since many urban expressways exist in the form of viaducts, when performing spatial aggregation of taxi GPS, in order to extract the actual traffic flow information reflected by realistic and effective in-service taxis, it is necessary to exclude vehicles running on urban roads under viaducts. Delete the GPS data of in-service taxis according to Figure 3 to determine the in-service taxis running on the urban expressway during the time period corresponding to each accident.

[0033] Determine the GPS trajectory data of all in-service taxis in each target time period of the target trajectory set as the second characteristic data. The GPS trajectory data for in-service taxis running on the urban expressway within the time range corresponding to each accident is the second characteristic data.

[0034] In actual application, since the GPS trajectory data of taxis can only obtain the GPS position of taxi driving and the instantaneous speed of the vehicle, only the characteristics of traffic flow speed are extracted based on the second characteristic data. Calculating the third characteristic data and the fourth characteristic data based on the second characteristic data is specifically It includes calculating the average speed of each in-service taxi in each road section of each viaduct of the second characteristic data based on the second characteristic data and the length of each road section of each viaduct. The calculation formula is as follows.

[0035]

Number

[0036] Here, n is the total number of GPS trajectory data of in-service taxi p in road section k of a specific viaduct, Vq represents the speed value of the q-th GPS trajectory data of in-service taxi p, L is the length of road section k of a specific viaduct, and as the third characteristic data, Vp indicates the average speed of in-service taxi p in road section k of a specific viaduct.

[0037] Based on the average speed of each operating taxi in each road section of each viaduct of the second characteristic data, calculate the average speed in each road section of each viaduct, and the calculation formula is as follows.

[0038]

Number

[0039] Here, M is the number of vehicles in road section k of a specific viaduct, and as the fourth characteristic data, V k is the average speed of road section k of a specific viaduct.

[0040] In actual application, based on the third characteristic data and the fourth characteristic data, using the D-S evidence theory to obtain the fusion average speed, specifically, The D-S evidence theory can construct an effective fusion model for heterogeneous sensor data, and its principle is as shown in Figure 6. The part (a) of Figure 6 represents the D-S data fusion method, that is, all data is unified and fused by the D-S data theory fusion method, and finally the fused data is output. The part (b) of Figure 6 shows another D-S data fusion method. Two sets of data are fused respectively, the fused data and another set of data are fused, and the fusion is carried out until all data is fused, and finally the fused data is output. The effects of these two methods are equivalent.

[0041] The specific steps are as follows. Step 1: Based on the third characteristic data and the fourth characteristic data, construct the following knowledge identification framework of the DS evidence theory.

[0042]

Number

[0043] V mircrove (t) and V GPS (t) are respectively derived from the fourth characteristic data and the third characteristic data, and V mircrove(t) represents the average speed of a specific road section based on cross-sectional microwave measuring instruments in the t time period, that is, the average speed of operating taxis in the road section of the viaduct, V GPS (t) represents the average speed of a specific road section based on taxi GPS trajectory data in the t time period, that is, the average speed of the road section on the viaduct. Since the two can be considered mutually exclusive, the power set composed of all elements of F(t) is as follows.

[0044]

Number

[0045] Here, X1(t) = V mircrove represents that the decision is the average speed of a specific road section of the viaduct of the fourth characteristic data, and X2(t) = V GPS represents that the decision is the average speed of an operating taxi in a specific road section of the viaduct of the third characteristic data, and X3(t) = V mircrove (t) ∩ V GPS (t) is an uncertain decision, indicating that it is impossible to distinguish which of the decisions of X1(t) and X2(t) it is.

[0046] Step 2: Under two types of data sources, determine the basic belief assignment for different decisions as follows. The assignment of the basic belief function is a quantitative criterion for the degree of support for a specific decision and is also evidence for each decision.

[0047]

Number

[0048]

Number

[0049] Here, i represents the evidence provided by the i-th data source, j represents the j-th decision, and m i (Xj (t) represents the degree to which the evidence provided by the i-th data source supports X j (t), and is called the basic distribution value of the decision X j (t), denoted as P j (X j (t)) is the basic probability assignment function of X j (t) in the i-th data source, where m i (φ) represents the basic distribution value of the decision φ, and X j (t) are the three types of decisions in Step 1.

[0050] P i (X j (t)) is assumed to satisfy V i (t) ~ N(u i (t), σ 2 i (t). Here, u i (t) and σ 2 i (t) represent the mean and variance of the average speed of the road section in the t time period in the historical data of the i-th type of data source, respectively. V i (t) represents the average speed of the road section in the historical data of the i-th type of data source in the t time period, and P i (X1(t)) represents the basic probability assignment function of X1(t) in the i-th data source, and P i (X2(t)) represents the basic probability assignment function of X1(t) in the i-th data source, and P i (X3(t)) represents the basic probability assignment function of X1(t) in the i-th data source.

[0051] P i (X j (t)) is calculated as follows.

[0052]

Equation

[0053] Step 3: Synthesize the two types of data sources as follows based on the Dempster evidence of the third characteristic data and the fourth characteristic data.

[0054]

Number

[0055] Here, K represents the contradiction between the evidences provided by the two types of data sources, m1(B) represents the basic probability assignment function of the third characteristic data, m2(C) represents the basic probability assignment function in the fourth characteristic data. The closer 1 / K is to 0, the greater the contradiction of the evidences provided by different data sources. A represents the above uncertain decision, B indicates that the above decision is the average speed of a specific road section in the third characteristic data, and C indicates that the above decision is the average speed of a specific road section in the fourth characteristic data.

[0056] Step 4: Determine the weights of the two types of data sources as follows according to the weight calculation formula of speed.

[0057] The weight calculation formula based on the average speed in the t time period of the third characteristic data is as follows.

[0058]

Number

[0059] The weight calculation formula based on the average speed in the t time period of the fourth characteristic data is as follows.

[0060]

Number

[0061] Finally, substituting the above formulas, the average speed in the t time period after fusing the two types of data sources as follows can be obtained.

[0062]

Number

[0063] Here, m(X1(t)) represents the basic distribution value of the fourth characteristic data source, m(X2(t)) represents the basic distribution value of the third characteristic data source, ωmircrove is the weight of the third characteristic data, and ωGPS is the weight 1-α(t) of the fourth characteristic data.

[0064] Since the traffic flow situation before an accident is closely related to the magnitude of the accident risk, it is called the pre-accident situation. And a traffic situation where an accident has not occurred under specific conditions is defined as a normal traffic flow state. In order to study and judge the change in traffic flow state between the time of accident occurrence and the normal traffic flow state, it is necessary to set up a control group (i.e., data of the normal traffic state) for comparison with the data of the accident group. According to the non-accident group selection step shown in Figure 4, the data of the control group is clarified. Since the occurrence of a traffic accident is a low-probability event, the experimental data of the control group is much larger than the experimental data of the accident group. Therefore, the creation of a real-time urban highway traffic accident risk prediction model becomes a problem of unbalanced data classification, that is, the number of samples in the majority class is much larger than the number of samples in the minority class. So far, in the creation of traffic accident prediction models, the "case-control" method has been mainly used for the problem of unbalanced data classification, that is, the data of the state where a traffic accident has occurred is selected as the case, and the corresponding data of the non-accident state is used as the control. In many subsequent studies, based on empirical methods, the ratio of the data of the case group to the control group is set to 1:4. However, if the ratio of 1:4 is selected based on empirical methods, a large amount of control group data will not be included in the model, and the prediction performance of the model may be impaired. Therefore, in the present invention, as a method for processing the unbalanced sample set, the Balanced Bagging integrated sampling method is adopted. Using Balanced Bagging to process the accident occurrence probability, the first characteristic data, the second characteristic data, the third characteristic data and the fourth characteristic data to obtain a training set. The principle of the Balanced Bagging algorithm is as shown in Figure 5. A set composed of control group data and experimental group data is directly divided into subsets that include both control group data and experimental group data. Resampling of control group data and experimental group data is performed in different subsets, and the traffic accident probability prediction model is classified to generate a plurality of classifiers. Next, the results of all classifiers are weighted and combined to obtain the final training set.

[0065] In actual application, training and optimizing an urban expressway traffic accident risk prediction model using a training set to obtain an optimized model specifically includes the following. Train an urban expressway traffic accident risk prediction model constructed using the GBDT algorithm using the training set to obtain a trained urban expressway traffic accident risk prediction model. Optimize the urban expressway traffic accident risk prediction model trained using the XGBoost algorithm to obtain an optimized model.

[0066] In actual application, the step of training an urban expressway traffic accident risk prediction model constructed using the GBDT algorithm using the training set to obtain a trained urban expressway traffic accident risk prediction model is specifically as follows.

[0067] Step 1: Construction of the initial accident risk prediction objective function y0':

[0068]

Number

[0069] y i represents the occurrence probability of a traffic accident corresponding to the i-th sample group of the training set. One sample group includes the occurrence probability of the corresponding traffic accident, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data. c represents the predicted value obtained by inputting the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data corresponding to the i-th data group into the accident risk prediction model. h it represents the decision tree model for the t-th iteration, and N represents the total number of sample groups.

[0070] The loss function can be defined as follows.

[0071] JPEG0007714263000015.jpg1491

[0072] Here, n represents the total number of all samples included in the training set, and y’ t represents the accident risk prediction objective function for the t-th iteration.

[0073] The accident risk prediction model can be defined as follows.

[0074]

Equation

[0075] Here, h it is the decision tree model after the t-th iteration.

[0076] Step 2: Train the model and perform iterations from t = 1 to T as follows. 1) Calculate the residual value r of the accident risk classification (i = 1~N) for each experimental group and control group at the t-th iteration, and y’ it represents the accident risk prediction objective function for the (t - 1)-th iteration. t-1 is represented by the accident risk prediction objective function of the (t - 1)-th iteration.

[0077]

Equation

[0078] 2) Use the training set to construct the decision tree model h it :

[0079]

Equation

[0080] 3) Minimize the loss function to calculate the weight β t and update the objective function y t ’. JPEG0007714263000019.jpg970JPEG0007714263000020.jpg1275JPEG0007714263000021.jpg2876

[0081] 4) Update the number of iterations t, proceed to the next iteration, return to 1), and repeat until t = T, then a trained urban expressway traffic accident risk prediction model is obtained.

[0082] Step 3: Output a trained traffic accident risk prediction model for urban expressways.

[0083] The traffic accident risk prediction model established in the urban expressway according to the present invention has good application prospects. The present invention enables early warning of accident risks based on traffic flow parameters, and traffic managers can timely transmit relevant traffic information in order to take appropriate safety measures.

[0084] The present invention realizes a cross-sectional microwave measuring instrument for measuring traffic flow data that falls into the full-sample detection category. However, what can be collected is only the traffic flow situation of the cross-section, and it is not possible to obtain the traffic flow information of the entire road section. The GPS trajectory data of taxis can effectively overcome this drawback by continuously tracking the traffic information of the entire road section. As shown in FIG. 1, the present invention proposes to use the theoretical method of DS evidence to integrate two types of data sources and provide a database for predicting traffic accident risks. The Balanced Bagging resampling method is used to balance the accident sample data and non-accident sample data, and more appropriately solve the unbalanced classification problem related to real-time accident prediction. Finally, the "black box model" of real-time accident risk prediction on urban expressways is interpreted based on feature importance analysis and partial dependence diagrams, and relevant analyses are performed.

[0085] According to the present invention, the traffic management department can early warn of the risk of traffic accidents, plan traffic police management in a timely manner, and more effectively improve the traffic safety level of the urban road network.

[0086] The present invention can effectively predict the accident occurrence rate on urban roads, enabling traffic-related management departments to accurately give early warnings of accident risks, timely create traffic management schedules, avoid traffic accidents, and effectively improve the traffic safety level of the urban road system.

[0087] Regarding model creation, the present invention is driven by the XGBoost algorithm, which is an integrated machine learning algorithm. The model created by the XGBoost algorithm can better fit the non-linear relationship between real-time traffic flow characteristics and accident risks on urban highways and has better interpretability.

[0088] In one embodiment, a computer device is further provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for predicting the occurrence probability of traffic accidents on urban expressways described in the embodiments of the above methods is realized. The computer device may be a database, and its internal structure diagram is as shown in FIG. 7. The computer device includes a processor, a memory, an input / output interface (abbreviated as Input / Output, I / O), and a communication interface. Here, the processor, the memory, and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. Here, the processor of the computer device is used to provide computing and control functions. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for executing the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store pending transactions. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, a type of data processing method is realized.

[0089] In one embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the above prediction method for the occurrence probability of traffic accidents on urban expressways described in the embodiments of the above method is realized.

[0090] It is necessary to explain that all the target information (including but not limited to target device information, target personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) related to this application are information and data that have been approved by the subject or fully approved by all relevant parties. Also, the collection, use, and processing of related data must comply with relevant laws, regulations, and standards.

[0091] The technical features of the above embodiments can be combined in any way. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered within the scope of this specification.

[0092] In this specification, specific examples are used to explain the principles and implementation methods of the present invention. The description of the above embodiments is only used to understand the method of the present invention and its central concept. At the same time, for those skilled in the art, based on the concept of the present invention, changes will occur in specific implementations and application scopes. In summary, the content of this specification should not be construed as limiting the present invention.

Claims

1. The "urban expressway" refers to a road targeted only on the viaduct, A method for predicting the occurrence probability of traffic accidents on an urban expressway, (A) A step of obtaining traffic flow information from each of a plurality of lanes of an urban expressway according to a predetermined collection frequency, and obtaining a set of traffic flow information for a plurality of lanes as a set of information flows; (B) For the traffic flow information for each lane included in the set of information flows, Testing the validity of the data collected by the measuring instrument to obtain traffic flow information having validity, and when there is discretely missing data in the traffic flow information determined to be valid, Estimating and complementing using the random walk method to obtain the estimated traffic flow information; And a step of obtaining the estimated traffic flow information; A step of determining such valid traffic flow information and the traffic flow information obtained by estimation and complementation as first characteristic data; (C) A step of collecting GPS trajectory data of all operating taxis in the city where the target road is located and accident records of traffic accidents occurring during the same time period within a predetermined set time period, and obtaining a GPS trajectory set and an accident record set; Here, the "set time period" refers to a time period having a predetermined time width before and after including the accident occurrence time described in the accident record set, and refers to a time range set for extracting the traffic flow related to the accident occurrence. (D) A step of determining second characteristic data by performing the following processing based on the accident record set and the GPS trajectory set, A step of extracting GPS trajectory data of taxis traveling on the urban expressway (on the viaduct) by deleting GPS trajectory data of taxis traveling on the urban road under the viaduct; Among the extracted data, a step of obtaining, as second characteristic data, GPS trajectory data of all operating taxis traveling on the urban expressway on the viaduct during the set time period corresponding to each accident record; (E) A step of calculating third characteristic data and fourth characteristic data according to the following procedure based on the second characteristic data, (E1) Referring to the second characteristic data after the under-viaduct data is deleted in the step (D), calculating the average speed of each operating taxi traveling on each road section of each viaduct constituting the urban expressway for each road section, and using the average speed per taxi as the third characteristic data; ​ Step (E2): Further aggregate the average speed per taxi obtained only from the road sections on the viaduct in step (E1) to calculate the average speed of the entire road section, and use this as the fourth characteristic data; Step (F): Integrate the third characteristic data and the fourth characteristic data using a data fusion method based on D-S evidence theory to obtain the fused average speed in the road section; Step (G): Input the fused average speed, the first characteristic data, the second characteristic data, the third characteristic data, and the fourth characteristic data into a real-time urban highway traffic accident risk prediction model to obtain the occurrence probability of a traffic accident; Step (H): Process the occurrence probability of the traffic accident and the first to fourth characteristic data using Balanced Bagging, which is a sampling method for class-imbalanced data, to obtain a training set; Step (I): After training an urban highway traffic accident risk prediction model constructed by the GBDT algorithm using the training set, optimize the model using the XGBoost algorithm to obtain an optimized model, and use the optimized model to predict the occurrence probability of traffic accidents on urban highways; A method for predicting the occurrence probability of traffic accidents on urban highways, characterized by including the above steps.

2. In the prediction method according to claim 1, The effectiveness test in step (B) is a process of determining whether the obtained traffic flow information contains abnormal values of sensors, inconsistencies in collection timing, and continuous deficiencies above a certain threshold. The traffic flow information determined to be normal in this test is obtained as "effective traffic flow information", and if there are partial deficiencies, they are estimated and complemented using the random walk method. A method for predicting the occurrence probability of traffic accidents on urban highways, characterized by this.

3. In the prediction method according to claim 1, In step (D), the reason for "deleting the GPS trajectory data of taxis running on the urban road under the viaduct" is to distinguish between the elevated part and the flat road of the urban highway and accurately extract the traffic flow situation on the viaduct. A method for predicting the occurrence probability of traffic accidents on urban highways, characterized by this.

4. In the prediction method according to claim 1, The "average speed of each operating taxi" in (E1) refers to the average speed of a taxi traveling on a specific road section of a specific viaduct, which is calculated from the length of the road section and the continuous GPS data of the taxi, and is a method for predicting the occurrence probability of traffic accidents on an urban expressway.

5. In the prediction method according to claim 1, The "average speed of each road section" in (E2) is the average speed of the entire road section obtained by aggregating the average speeds of a plurality of taxis calculated in (E1), and is a method for predicting the occurrence probability of traffic accidents on an urban expressway.

6. In the prediction method according to claim 1, The D-S evidence theory in (F) is a probabilistic inference method considering uncertainty based on the Dempster-Shafer theory, and is a data fusion method for calculating a fused average speed by integrating evidence obtained from a plurality of different data sources (such as the third characteristic data and the fourth characteristic data), and is a method for predicting the occurrence probability of traffic accidents on an urban expressway.

7. In the prediction method according to claim 1, Balanced Bagging in (H) is a method for extracting samples of both classes in a balanced manner in order to reduce the imbalance between accident data of the minority class and non-accident data of the majority class, and is a method for generating a final training set by creating and resampling a plurality of subsets and then combining the results of a plurality of classifiers trained with those subsets, and is a method for predicting the occurrence probability of traffic accidents on an urban expressway.

8. A computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for predicting the occurrence probability of traffic accidents on an urban expressway according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Accident prediction method, computer program, accident prediction device, and learning model generation method

    JP2021182189A

  • Traffic control device and learning model production method

    JP2022074223A

  • Information processing device

    JP2023086112A

Cited By

  • Cooperative control method for vehicles in diverging and converging areas

    CN120877527A