Data-driven path selection and dynamic feedback expressway traffic volume prediction method

By using multi-source data fusion and the path selection and dynamic feedback method of the LightGBM model, the problem of insufficient accuracy in highway traffic volume prediction is solved, and high-precision, real-time traffic flow state prediction is achieved, providing reliable data support for traffic planning and management.

CN121600740APending Publication Date: 2026-03-03GUANGZHOU TRANSPORTATION RES INST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511871998.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing methods for predicting highway traffic volume are insufficient to reveal the intrinsic relationship between traffic flow growth and urban industrial activities and spatial structure evolution in cross-regional and long-distance transportation links. Furthermore, traditional methods suffer from insufficient prediction accuracy, model generalization ability, and real-time performance.

Method used

By employing multi-source data fusion and governance, a three-dimensional feature vector system is constructed. The LightGBM model is used to calculate the path selection probability, and dynamic traffic allocation is achieved through iterative allocation. The LightGBM machine learning model is used to predict the path selection probability and road segment travel time, and the road network status is dynamically fed back.

Benefits of technology

It achieves high-precision, real-time route selection and traffic flow state prediction, provides reliable basic data to support traffic planning and management, and improves the accuracy and adaptability of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600740A_ABST
    Figure CN121600740A_ABST
Patent Text Reader

Abstract

The invention discloses a data-driven path selection and dynamic feedback expressway traffic volume prediction method, which comprises the following steps: S1, multi-source data fusion and treatment, S2, multi-model fusion path selection probability calculation, S3, dynamic traffic distribution based on road resistance iteration, S4, data-driven road resistance calculation and dynamic feedback, and S5, data-driven path selection and dynamic feedback. According to the invention, on the basis of fusion analysis of expressway toll station, toll portal running water data and weather data, traffic flow information, weather information and the like running on an expressway network can be comprehensively obtained, and feature construction of all factors influencing path selection is realized; lightGBM is selected to establish a path selection probability prediction model, and real-time prediction of the path selection probability of each trip is realized; and finally, in combination with a flow iteration dynamic distribution model, charging OD data is dynamically restored to a road-section-level and high-time-resolution road-section traffic flow state with high precision, and reliable basic data is provided for traffic planning, engineering project early-stage decision making, emergency management, real-time induction and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of traffic engineering and intelligent transportation technology, specifically relating to a data-driven method for route selection and dynamic feedback for predicting highway traffic volume. Background Technology

[0002] Traffic volume forecasting is a core component of feasibility studies, operation management, and traffic control strategy formulation for highway construction projects. Its forecasting accuracy directly determines the determination of project technical level, engineering scale, economic evaluation results, and the effectiveness of traffic management measures. Currently, my country's highway traffic volume forecasting mainly adopts the four-stage method derived from urban traffic planning. This method completes the forecasting through four stages: trip generation, trip distribution, mode classification, and traffic allocation. However, it has significant limitations in the highway scenario.

[0003] Existing highways, with their core function of cross-regional, long-distance transportation connections, rely heavily on existing four-stage methods that simulate the impact of socio-economic factors on traffic flow growth at a macro-regional level. These methods fail to reveal the intrinsic connection between traffic flow growth and urban industrial activities and spatial structure evolution, and neglect the functional positioning and traffic characteristics of specific construction projects, resulting in insufficient prediction accuracy. Currently, many highway projects suffer from decision-making misguidance and implementation risks due to traffic volume prediction errors. Given increasing investment and financing pressures, there is an urgent need for prediction methods more closely aligned with the characteristics of highway systems. To improve prediction accuracy, domestic and international scholars have attempted to apply data-driven methods such as machine learning to route selection modeling and traffic flow prediction. YAMAMOTO et al. used GPS trajectory data and employed decision trees and neural networks to analyze driver route selection behavior; however, these models are limited by data scale and computing power. The generalization ability and real-time performance are insufficient. PARK et al. proposed the fuzzy ID3 algorithm to solve the information loss problem of continuous variable discretization, but the setting of fuzzy rules depends on expert prior knowledge, which is highly subjective and not conducive to large-scale promotion. SUN and PARK used support vector machines to build a path selection prediction model. Although the accuracy is higher than that of traditional models, it has the characteristics of a "black box" and the decision-making process lacks interpretability. LAI et al. compared a variety of machine learning models and proved that the ensemble learning method has the best performance, but it still has the problems of poor interpretability and not involving the generation of alternative path sets. SIFRINGER et al. fused representation learning and discrete selection models, which improved the prediction accuracy, but the model structure is complex and the latent representation is difficult to interpret. YAO and BEKHOR proposed a data-driven two-stage framework to generate alternative path sets, but its effectiveness depends on the quality and coverage of trajectory data. Summary of the Invention

[0004] The purpose of this invention is to provide a data-driven route selection and dynamic feedback method for predicting highway traffic volume, in order to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a data-driven route selection and dynamic feedback method for predicting highway traffic volume, comprising the following steps:

[0006] S1. Multi-source data fusion and governance: Collect traffic flow data from highway toll stations, gantry traffic flow data, and weather data from meteorological stations as basic data sources. Then, perform abnormal and redundant data removal, missing data supplementation, spatiotemporal correlation fusion of multi-source data, and feature engineering processing on the collected data in sequence to construct a three-dimensional feature vector system that includes individual features, traffic flow operation features, and weather features.

[0007] S2. Path selection probability calculation using multi-model fusion: for all highway trips (OD stands for Origin-Destination). Construct a path selection feature vector set, where m is the total number of highway trip ODs. For each trip, use the shortest path algorithm to search for k valid paths. The feature vector set for each path is calculated as follows:

[0008]

[0009] in, , ,... These represent the k valid paths found using the shortest path algorithm;

[0010] This represents the first feature value of the first path. This represents the third feature value of the second path. The total number of features is represented by a matrix, which systematically represents the path and features of each trip. The m×k path feature vectors from m trips are input into the LightGBM model for training, establishing the relationship between traffic volume and traffic flow features, and k-fold cross-validation is used. The optimal parameters of the LightGBM model are obtained and used to characterize the mapping relationship between path-related features and path selection probabilities:

[0011] ;

[0012] S3. Dynamic flow allocation based on road resistance iteration: setting initial travel time for all road segments e.

[0013] (Free-flow time) Initialize the load of all road segments in each allocated time slice T. The dynamic capacity value for all road segments is set to the design capacity. The system iterates through the allocation process, calculating the predicted travel time for each path based on the road network status after the previous allocation (after the (n-1)th allocation). And calculate the predicted travel time using path k. Use the calculated path time Update the feature vectors of each path, then input them into the pre-trained path selection model to obtain the selection probability of each path. And the traffic demand of the OD. According to probability Traffic is allocated to each path, resulting in incremental traffic distribution. ;

[0014] in : indicates all road segments Set initial travel time Its value is equal to the free-flow time. Free-flow time refers to the time it takes for vehicles to pass through a road segment when traffic volume is very low and there is no road congestion. The time required;

[0015] Initialize the load of all road segments to 0 in each allocated time slice T. The load here can be understood as the load of the road segments within a specific time slice T. The amount related to the traffic pressure is initially set to 0 to indicate that there is no traffic load at the beginning;

[0016] This represents the dynamic capacity value of all road segments. It is set to the design capacity. Capacity refers to the maximum number of vehicles or pedestrians that a traffic facility can pass through per unit time under certain traffic conditions. Dynamic capacity takes into account factors such as changes in traffic conditions over time.

[0017] Based on the previous round of allocation (the 1st round) After the first allocation, the predicted travel time for each path is calculated based on the road network status. As the allocation iterates, the road network status changes after each allocation (e.g., changes in road segment traffic flow), thus affecting the travel time of road segments. This is based on the first... The travel time for each segment in the next allocation is based on the previous ( Prediction of the next state;

[0018] : Use path The predicted travel time. It takes into account the entire route. The predicted travel time for each road segment is calculated based on factors such as the predicted travel time for each segment, and is used to measure the speed of travel in the first segment. During the next allocation, the path is used. Time required;

[0019] The probabilities of choosing each path are obtained through a pre-trained path selection model. In traffic assignment, travelers have different preferences for different paths, and these probabilities reflect the choices made on different paths. Path selection during secondary allocation The probability of;

[0020] Traffic demand at the origin-destination (OD); Indicates the first In the secondary allocation, the total traffic demand from a certain starting point to a certain ending point;

[0021] S4. Data-driven path resistance calculation and dynamic feedback: This involves calculating the path flow rate from S3. This is converted into a virtual vehicle trajectory for the road segments traveled. Based on the predicted travel time of each segment e on path k, the specific time window for each vehicle to enter and leave the segment is calculated, and the traffic flow of different vehicle types is converted into the equivalent of a standard passenger car according to the standard vehicle conversion factor, and accumulated into the corresponding segment-time slice load matrix: For road segment e that requires impedance updates, a feature vector is constructed in real time during time slice T. and the feature vector Input pre-trained LightGBM machine learning model Directly predict the travel time of this road segment under current conditions:

[0022] ;

[0023] The network impedance status is updated and fed back to the next round of allocation. Once all OD pairs have been allocated...

[0024] Then, output the segment-level road network traffic results with a time resolution of 5 minutes;

[0025] in, Indicates road segment In the The trip time at the time of calculation or update is...

[0026] This refers to the time required for a vehicle to travel through this section of road;

[0027] Preferably, in step S1, the toll station transaction data includes license plate number, vehicle type, entry / exit identifier, passage time, and payment type; the gantry transaction data includes license plate number, vehicle type, gantry unique identifier, passage time, and passage medium type; and the weather data includes temperature, relative humidity, average wind speed, and cumulative rainfall. This clarifies the specific data composition of each basic data source in the multi-source data fusion and governance steps, providing accurate data basis for subsequent feature extraction and model training.

[0028] Preferably, in step S1, the feature engineering process specifically involves: encoding the fused data, where individual features include the frequency of high-speed travel of vehicles, vehicle type, local administrative region, and energy type; traffic flow operation features include the difference in travel time, mileage, toll cost, and number of switching times between highways and ordinary roads between alternative routes and the shortest path; and weather features include the average temperature, average relative humidity, average wind speed, and cumulative rainfall for the corresponding time interval and road segment. After one-hot encoding or numerical processing, a 10-15 dimensional feature vector is formed, refining the feature engineering processing method and the specific composition of the three-dimensional feature vector. Standardized feature vectors are formed through encoding processing, providing high-quality input for model training.

[0029] Preferably, in S3, during the iterative allocation process, a road network impedance update is performed every 100-200 OD pairs allocated. When the load change rate of the road segment in two adjacent iterations is less than 5%, the iterative allocation of the OD pair can be terminated in advance. This clarifies the impedance update frequency and iteration termination conditions in dynamic flow allocation, improving allocation efficiency while ensuring prediction accuracy.

[0030] Preferably, in S4, the standard vehicle conversion coefficient is as follows: 1.0 for Type I passenger cars and Type I freight cars; 1.2 for Type II passenger cars and Type II freight cars; 1.5 for Type III passenger cars and Type III freight cars; 2.0 for Type IV passenger cars and Type IV freight cars; and 2.5-4.0 for Type V and above freight cars and special-purpose vehicles. The specific values ​​are determined according to the number of axles and load capacity of the vehicle. The standard vehicle conversion coefficients for different vehicle types are specified to ensure the accuracy of traffic flow conversion for different vehicle types when calculating road load and to improve the accuracy of road resistance prediction.

[0031] Preferably, in S4, the impedance feature vector includes six types of sub-features: flow characteristics including road segment-time slice load, upstream road segment flow, and downstream road segment flow; saturation characteristics including the ratio of road segment load to dynamic capacity; traffic composition characteristics including the flow proportion of each vehicle type; spatiotemporal characteristics including time interval identifier, road segment location identifier, and weekday attribute; dynamic event characteristics including the presence of traffic event identifier and event level; and historical state characteristics including the average road segment load of the previous three adjacent time slices. This clarifies the composition of the six types of sub-features of the impedance feature vector in road resistance calculation, providing comprehensive feature support for the LightGBM model to accurately predict road segment travel time.

[0032] Preferably, in step S4, the training samples of the pre-trained LightGBM machine learning model are historical road segment-time slice load data, corresponding impedance feature vectors, and actual travel time data from the past 3-6 months. The mean absolute error (MAE) of the model test set does not exceed 1.5 minutes. The source of training samples and accuracy requirements of the LightGBM model used for road resistance prediction are specified to ensure the reliability of the model's prediction of travel time.

[0033] Preferably, the model update step is also included: every 7-15 days, the LightGBM path selection model and the LightGBM regression impedance model are incrementally trained using the latest collected multi-source data to update the model parameters to ensure prediction accuracy. By adding the step of regular incremental model updates, the model performance is maintained by incorporating new data, ensuring the stability of long-term prediction accuracy.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] (1) Based on the fusion analysis of highway toll station, toll gantry flow data and weather data, this invention can comprehensively obtain traffic flow information and weather information on the highway network, and realize the feature construction of all factors affecting path selection. Considering accuracy and algorithm complexity, LightGBM (Light Gradient Boosting Machine) is selected to establish a path selection probability prediction model to realize real-time prediction of the path selection probability for each trip. Finally, combined with the flow iteration dynamic allocation model, the toll OD data is dynamically restored to the road segment level and high time resolution road segment traffic flow status with high precision, providing reliable basic data for traffic planning, early decision-making of engineering projects, emergency management, real-time guidance and other purposes.

[0036] (2) The data is easy to obtain, the sample covers a wide range of time and space, the analysis accuracy is high, it is data-driven, the algorithm is simple and feasible, and it has strong universality. It uses LightGBM to predict the path selection probability and is easy to put into practical application.

[0037] (3) By adopting a dynamic traffic allocation model, the traditional, fixed-parameter road resistance function is upgraded to a "data-driven dynamic impedance estimator" based on machine learning. This enables the learning of complex nonlinear mapping relationships between multiple factors such as traffic flow, saturation, vehicle type composition, weather, and events and travel time from historical data. This allows for a more accurate and adaptive reflection of the real road network status, demonstrating the impact of changes in road network operation status on vehicle selection and solving the problem that traditional dynamic allocation cannot simulate real path selection. Attached Figure Description

[0038] Figure 1 This is a flowchart of the method of the present invention;

[0039] Figure 2 This is a comparison diagram of the present invention with the traditional "four-stage method" and the "four-stage method" that considers traffic induction and diversion. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installed," "equipped with," "sleeved with," "connected," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances.

[0042] This invention provides, for example Figure 1-2 This illustrates a data-driven route selection and dynamic feedback method for predicting highway traffic volume.

[0043] Example

[0044] This invention is based on massive amounts of multi-source big data, including highway toll station data, gantry flow data, and weather data. First, it employs a data cleaning and preprocessing process to filter out erroneous and abnormal data, complete missing data, and correlate data based on spatiotemporal information. Second, it uses multi-model fusion to construct a vehicle travel choice model based on multiple factors such as mileage, cost, time, and weather, achieving more realistic path selection probability calculations. Finally, it constructs an iterative dynamic traffic allocation mechanism based on data-driven road resistance, simulating dynamic user equilibrium of traffic flow, addressing the correlation between flow and road resistance, and achieving high-precision dynamic reconstruction of toll OD data into road network-level, high-temporal-resolution segment traffic flow status. This provides reliable basic data for traffic planning, early-stage project decision-making, emergency management, and real-time guidance.

[0045] 1.1 Specific steps:

[0046] 1.1.1 Data Fusion and Governance

[0047] To achieve accurate traffic flow prediction on highways, this invention primarily uses highway traffic flow data (including toll station and ETC gantry traffic flow data) and weather data as foundational data for fusion analysis. Toll station data mainly includes license plate number, vehicle type, entry / exit information, passage time, and payment type; gantry data mainly includes license plate number, vehicle type, gantry section, passage time, and medium of transport; weather data includes temperature, humidity, wind speed, and rainfall. Because the above data sources are diverse and the sampling granularity varies, data fusion and processing are necessary to support the subsequent construction of accident risk prediction models.

[0048] Step 1: Remove abnormal and redundant data:

[0049] Toll station transaction data: Abnormal data such as "entry time is later than exit time" and "travel time exceeds the reasonable multiple of the maximum theoretical time of the OD" are removed; redundant data such as "the same license plate number enters and exits the same toll station continuously within a short period of time" are identified by the DBSCAN clustering algorithm, which significantly improves the cleanliness of the data;

[0050] ETC gantry flow data: Based on road network topology, abnormal data that "violates the road network connectivity in the order of vehicles passing through the gantry" is removed; reasonable speeding thresholds are set to remove abnormal speeding data, effectively handling abnormal data;

[0051] Weather data: Extreme outliers are removed using the 3σ criterion. Unreasonable records are directly marked as outliers and removed. A small number of outlier data are processed.

[0052] Step 2: Data Imputation. For missing weather data: For short-term missing data at a single station, the "spatiotemporal interpolation method" is used. The missing station is the center, and data from surrounding neighboring stations are selected. The imputation value is calculated by combining distance weights. For long-term missing data, the LSTM model is used to predict and impute based on historical data from the same period, which has high imputation accuracy.

[0053] Missing gantry data: When a small number of gantry records are missing between adjacent gantry, the data is supplemented based on the shortest path principle of the road network topology; when a large number of gantry records are missing, the data is supplemented by combining the toll station entry and exit times and the path preference model, resulting in a high accuracy rate.

[0054] Missing vehicle attribute data: For records with "missing vehicle type", the data is inferred by reverse deduction using the license plate prefix and toll standard; for records with "missing energy type", the data is completed by matching the charging facility usage records at toll stations with the vehicle brand and model database, with a high accuracy rate.

[0055] Step 3: Spatiotemporal correlation and fusion of multi-source data. First, the time field is processed. Due to the different collection frequencies of various data, all data are resampled at 5-minute intervals. Specifically, for highway toll data, highway gantry data, and weather data, if the time field is "December 2, 2024, 14:04", it is converted to a "December 2, 2024, 14:00 – December 2, 2024, 14:05" field (period). Second, the spatial field is processed. Since the spatial scale of highway toll data and highway gantry data is significantly larger than that of weather data recorded using station numbers, the station numbers are classified according to spatial location and by gantry segment (link). Finally, spatiotemporal correlation and fusion are performed based on the spatiotemporally processed period and link fields of the multi-source data.

[0056] Table 1 Example of Billing Transaction Information

[0057]

[0058] Table 2 Example of gantry flow data

[0059]

[0060] Table 3. Example of weather data

[0061]

[0062] After data fusion and governance are completed, the data is re-encoded according to data types such as discrete and continuous, and individual characteristics (including highway preference, vehicle type, geographical distribution, and vehicle energy type, mainly extracted from highway toll revenue and gantry revenue data), traffic flow characteristics (including travel time difference, travel mileage difference, travel cost difference, number of times of switching between primary highways, etc., mainly extracted from highway toll revenue and gantry revenue data), and weather characteristics (including wind speed, rainfall, humidity, and temperature).

[0063] Table 4. Schematic diagram of highway feature vector construction

[0064]

[0065] 1.1.2 Path selection probability calculation for multi-model fusion

[0066] After completing data fusion and governance, in order to scientifically and accurately predict the probability of path selection, this invention proposes to construct a path selection prediction model using LightGBM (Light Gradient Boosting Machine). This model predicts the probability of path selection by inputting real-time traffic flow characteristics and weather features. The specific approach is as follows:

[0067] 1. For all highway travel Construct a path selection feature vector set, where m is the total number of origin-destination (OD) trips on the highway, for each trip... The shortest path algorithm is used to search for k valid paths. The feature vector set for each path is calculated as follows:

[0068]

[0069] in, , ,... These represent the k valid paths found using the shortest path algorithm;

[0070] This represents the first feature value of the first path. This represents the third feature value of the second path. The total number of features is represented in this matrix form, which systematically represents the path and features of each trip.

[0071] 2. Input the m×k path feature vectors from m trips into LightGBM for training, and establish the relationship between traffic volume and traffic flow features, where... The optimal parameters of the LightGBM model are obtained through k-fold cross-validation, which are used to characterize the mapping relationship between path-related features and path selection probabilities.

[0072]

[0073] in, It is the set of adjustable parameters in the LightGBM model (such as learning rate, tree depth and other hyperparameters), while k-fold cross-validation is a method for model evaluation and parameter tuning. It divides the dataset into k parts, and alternately uses k-1 parts as the training set and 1 part as the test set, and performs k training and testing cycles in this way to find the optimal model parameters.

[0074] 1.1.3 Dynamic Flow Allocation Based on Road Resistance Iteration

[0075] For each trip, OD (Original Distance) is predicted and allocated in chronological order. After each allocation, the road network segment operation status is updated. Through iteration, a dynamic balance between traffic flow and road resistance is achieved, as detailed below:

[0076] (1) Step 1: Initialization

[0077] 1. Set the initial travel time for all road segments e. (Free-flow time), where :

[0078] Represents all road segments Set initial travel time Its value is equal to the free-flow time. Free-flow time refers to the time it takes for vehicles to pass through a road segment when traffic volume is very low and there is no road congestion. The required time is calculated based on the design speed and length of the road segment, using the formula "Free-flow time = Road segment length / Design speed × 60" (unit: minutes). For special road segments such as bridges and tunnels, historical free-flow data is multiplied by a correction factor to ensure it conforms to the actual situation.

[0079] 2. Initialize the load of all road segments in each allocated time slice T. ,in, :

[0080] The load of all road segments is initialized to 0 in each allocated time slice T. Here, the load can be understood as the load of the road segment within a specific time slice T. The amount related to the traffic pressure is initially set to 0 to indicate that there is no traffic load at the beginning. The unit of load is "standard passenger car equivalent (pcu)".

[0081] 3. Set the dynamic capacity value of all road segments to the design capacity. ;

[0082] (2) Step 2: Iteration allocation (Iteration n = 1 to m, where m is the total number of highway trip origins and destinations):

[0083] A. Dynamic segment time calculation: Based on the travel time of all highway network segments updated in the previous round (after the (n-1)th allocation). Calculate the predicted travel time using path k. This includes the real-time impedance based on the estimated arrival time at each road segment;

[0084] B. Traffic allocation based on path selection model:

[0085] 1. Use the calculated path time Update the feature vectors of each path, and then input the pre-defined path.

[0086] The trained path selection model yields the selection probability of each path. Given the updated feature vector of the journey time, the model outputs the probability as follows: the main path has the highest probability, followed by the parallel path, and the transfer path has the lowest probability.

[0087] 2. Traffic demand for the OD. According to probability Traffic is allocated to each path (primary path traffic allocation)

[0088] The most frequent path is parallel path, followed by transfer path, which has the least frequent path, resulting in incremental traffic allocation. .

[0089] C. Spatiotemporal reconstruction and load accumulation:

[0090] 1. Path traffic This is converted into a virtual vehicle trajectory for the road segments traveled. ,

[0091] Based on the predicted travel time of each segment e on path k, calculate the specific time window for each vehicle to enter and leave the segment;

[0092] 2. The airflow of different vehicle types is converted to the equivalent of a standard passenger car using the PCU coefficient, as detailed below:

[0093] Table 5 Standard Vehicle Conversion Factors

[0094] Vehicle type Standard vehicle conversion factor {1,2,11,21} 1 {3,4,12,22} 1.5 {13,23,14,24} 2.5 {15,16,25,26} 4

[0095] Vehicle types: 1-Type I passenger vehicle, 2-Type II passenger vehicle, 3-Type III passenger vehicle, 4-Type IV passenger vehicle, 11-Type I freight vehicle, 12-Type II freight vehicle, 13-Type III freight vehicle, 14-Type IV freight vehicle, 15-Type V freight vehicle, 16-Type VI freight vehicle, 21-Type I special-purpose vehicle, 22-Type II special-purpose vehicle, 23-Type III special-purpose vehicle, 24-Type IV special-purpose vehicle, 25-Type V special-purpose vehicle, 26-Type VI special-purpose vehicle;

[0096] 3. The converted equivalent flow rate is accumulated into the corresponding segment-time-slot load matrix according to the time slot occupied by segment e: This enables precise location of traffic on the spatiotemporal network. By converting the traffic of each path into road segment load, the road segment through which the main path passes increases the load the most, and the updated road segment load is the corresponding value.

[0097] D. Data-driven path resistance calculation and dynamic feedback:

[0098] 1. Feature Vector Construction: For the road segment e whose impedance needs to be updated, the feature vector is constructed in real time at time slice T. It includes the following features:

[0099] 1.1 Flow characteristics: Current equivalent load Upstream inflow, downstream outflow;

[0100] 1.2 Saturation characteristics: The dynamic traffic capacity is dynamically adjusted based on real-time road attributes (such as construction, lane closures, etc.);

[0101] 1.3 Traffic composition characteristics: proportion of small trucks, proportion of large trucks, etc.;

[0102] 1.4 Spatiotemporal characteristics: Time period (peak / off-peak), weekday / weekend;

[0103] 1.5 Characteristics of dynamic events: construction, weather, and accident situations;

[0104] 1.6 Historical status characteristics: Average speed and load of the previous two time slices (default is 5 minutes);

[0105] 2. Travel time prediction: The feature vector... Input pre-trained LightGBM machine learning model It directly predicts the travel time of this road segment under current conditions:

[0106]

[0107] 2.1 Iterate and output repeatedly until all ODs are assigned, and output the final road network traffic allocation result;

[0108] To ensure the long-term prediction accuracy of the model, a mechanism of "regular incremental training + real-time monitoring and adjustment" is established:

[0109] Regular incremental training: Incremental training of the model is carried out every 10 days. The latest 10 days of multi-source data are used as incremental samples. After being fused with historical samples, the model parameters are updated. Incremental training only requires retraining the decision tree nodes corresponding to the newly added samples. The training time is controlled within 1 hour to avoid the time-consuming problem of full training.

[0110] Real-time monitoring and adjustment: Establish a model accuracy monitoring dashboard to calculate the error between predicted and actual values ​​in real time. When the MAE (mean absolute error) of a certain road segment is greater than 2.0 minutes for three consecutive time slices, an emergency adjustment process is triggered. The real-time data of the road segment in the most recent hour is used for rapid fine-tuning to ensure that the model accuracy is restored to a qualified level.

[0111] The present invention proposes a data-driven path selection and dynamic feedback method for predicting highway traffic volume. The path enumeration method in the prediction method adopts the k-shortest path method. In practical applications, depending on the availability of actual data or computational performance, a breadth-first search with limited number of searches or a restriction based on path similarity can be adopted to more effectively generate the corresponding effective path set.

[0112] Alternatives to data sources and feature engineering: In addition to toll station and gantry traffic data, GPS /

[0113] Beidou floating vehicle.

[0114] Trajectory data or mobile signaling data can be used to obtain more continuous and comprehensive travel time and route information, which is especially suitable for non-ETC vehicles or route reconstruction.

[0115] An alternative to path selection prediction models: LightGBM as an efficient gradient boosting tree model is

[0116] While XGBoost, CatBoost, Random Forest, and Deep Neural Networks are preferred, they can all be used as alternative models when dealing with extremely complex nonlinear relationships.

[0117] In addition to using LightGBM regression to directly predict travel time, data-driven road resistance calculation and dynamic feedback can also consider using deep learning methods (such as spatiotemporal graph neural networks) to directly predict and update the road network status, thus capturing the spatiotemporal correlation between road segments more naturally.

[0118] And such as Figure 2As shown, the "four-stage method" is commonly used for road traffic volume forecasting both domestically and internationally. The traditional "four-stage" forecasting method is based on a person trip survey and consists of four stages: Trip Generation, Trip Distribution, Model Split, and Trip Assignment. The ultimate goal of the "four-stage method" is to provide a scientific basis for regional road network planning. Its general steps are: first, based on the regional economic growth trend, obtain raw data through current regional socio-economic surveys and traffic surveys; second, use forecasting methods and models to deduce the overall future incremental travel demand of the region; third, predict the distribution of future traffic generation and attraction across regions to obtain the total inter-regional traffic travel demand; fourth, obtain the traffic volume carried by each mode of transportation through model allocation; and finally, rationally allocate all travel demand on the road network, while simultaneously verifying the load of the existing road network in conjunction with road network planning.

[0119] Compared to the traditional "four-stage method" used in urban transportation planning, the traffic volume prediction method in the context of highways combines its actual operational characteristics and needs, dividing traffic generation into three parts: trend-based, transfer-based, and induced-increase traffic volume. This method comprehensively considers the traffic induced-increase effect brought about by future road network construction and land development, as well as the transfer impact between other modes of transportation, thereby improving the accuracy of traffic volume prediction. In addition, since the transportation modes in the highway system are singular, the traffic mode division stage is usually omitted. Therefore, the "four-stage" model is simplified to three links: traffic generation, traffic distribution, and traffic assignment. The key link is to accurately reconstruct the driving path of vehicles in the actual road network (traffic assignment) based on the OD data (traffic distribution) of the highway toll station entrances and exits, that is, to establish an effective path selection model.

[0120] Route selection models aim to describe travelers' route selection behavior from origin to destination, falling under the category of disaggregate behavior modeling. The core task of this modeling is to quantitatively analyze how travelers choose routes based on reasonable behavioral assumptions, and to estimate and predict their actual driving conditions within the traffic network. In recent years, with the development of holographic traffic perception technology, massive amounts of vehicle trajectory data have provided strong support for route selection behavior research, driving significant progress in this field. Previous studies largely relied on labeled questionnaire data, which had limitations such as limited sample size and high survey costs. Advances in data collection technology have provided richer spatiotemporal movement data and individual attribute information, making machine learning-based methods possible in route selection modeling.

[0121] Machine learning, as a data-driven modeling paradigm, automatically learns its model structure from data through algorithms, rather than relying on researchers' pre-defined structures. Compared with traditional discrete choice models driven by theory, machine learning methods have the following advantages: First, they avoid the difficulties in optimizing model structures and the biases that may be caused by incorrect structure settings in traditional methods; second, they usually perform better in terms of goodness of fit; third, they are suitable for training on large-scale datasets, which better meets the modeling needs in big data environments; and fourth, they can handle special types of data such as text, images, and continuous data streams, thereby expanding the research dimensions of path selection modeling.

[0122] References in the background section (such as patents / papers / journals):

[0123] 1.YAMAMOTO T,KITAMURA R,FUJ J. DriversRoute Choice Behavior: Analysis by Data Mining Al-gorithms [J]. Transportation Research Record, 2002(1807):59-66.

[0124] 2.PARK K, BELL MGH, KAPARIAS I, et al. SoftDiscretization in aClassification Model for ModelingAdaptive Route Choice with a Fuzzy ID3Algorithm[J].Transportation Research Record, 2008(2076):20-28.

[0125] 3. SUN B, PARK BB. Route Choice Modeling withSupport Vector Machine[J]. Transportation Re-search Procedia, 2017, 25:1806-1814.

[0126] 4.LAIX,FU H, Ll, et al. Understanding DriversRoute Choice Behaviors in the Urban Network withMachine Learning Models [J]. IET Intelligent Trans-port Systems, 2018,13(3):427-434.

[0127] 5. Wang Luyao, Jiang Xi. Modeling of passenger route selection in urban rail transit based on ensemble learning [J]. Journal of Railway Engineering, 2020, 42(6):18-24.

[0128] 6.SIFRINGER B, LURKIN V,ALAHI A. Enhancing Discrete Choice Models withRepresentation Learning[J]. Transportation Research Part B, 2020, 140:236-261.

[0129] 7. YAO R, BEKHOR S. Data-driven Choice Set Gener-ation and Estimation of Route Choice Models [J].Transportation Research PartC, 2020.121:102832.

[0130] 8.Meng Q .LightGBM: A Highly Efficient Gradient Boosting DecisionTree[C] / / Neural Information Processing Systems.Curran Associates Inc. 2017.

[0131] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A data-driven route selection and dynamic feedback method for predicting highway traffic volume, characterized in that, Includes the following steps: S1. Multi-source data fusion and governance: Collect traffic flow data from highway toll stations, gantry traffic flow data, and weather data from meteorological stations as basic data sources. Then, perform abnormal and redundant data removal, missing data supplementation, spatiotemporal correlation fusion of multi-source data, and feature engineering processing on the collected data in sequence to construct a three-dimensional feature vector system that includes individual features, traffic flow operation features, and weather features. S2. Path selection probability calculation using multi-model fusion: for all highway trips Construct a path selection feature vector set, where m is the total number of origin-destination (OD) routes on the highway. For each trip, use the shortest path algorithm to search for k valid paths. The feature vector set for each path is calculated as follows: , The feature vectors of m×k routes from m trips are input into the LightGBM model for training to establish the relationship between traffic volume and traffic flow features, and k-fold cross-validation is used. The optimal parameters of the LightGBM model are obtained and used to characterize the mapping relationship between path-related features and path selection probabilities: , in, It is the set of adjustable parameters in the LightGBM model; S3. Dynamic flow allocation based on road resistance iteration: setting initial travel time for all road segments e. Initialize the load of all road segments in each allocated time slice T. The dynamic capacity value for all road segments is set to the design capacity. The system iterates through the allocation process, calculating the predicted travel time for each path based on the road network status after the previous allocation. And calculate the predicted travel time using path k. Use the calculated path time Update the feature vectors of each path, then input them into the pre-trained path selection model to obtain the selection probability of each path. And the traffic demand of the OD. According to probability Traffic is allocated to each path, resulting in incremental traffic distribution. ; S4. Data-driven path resistance calculation and dynamic feedback: This involves calculating the path flow rate from S3. Transformation For virtual vehicle trajectories, the road segments traveled... Based on the predicted travel time of each segment e on path k, the specific time window for each vehicle to enter and leave the segment is calculated, and the traffic flow of different vehicle types is converted into the equivalent of a standard passenger car according to the standard vehicle conversion factor, and accumulated into the corresponding segment-time slice load matrix: For road segment e that requires impedance updates, a feature vector is constructed in real time during time slice T. and the feature vector Input pre-trained LightGBM machine learning model Directly predict the travel time of this road segment under current conditions: ; It updates the network impedance status and feeds it back to the next round of allocation. After all OD pairs are allocated, it outputs the segment-level network traffic results with a time resolution of 5 minutes.

2. The data-driven route selection and dynamic feedback method for predicting highway traffic volume according to claim 1, characterized in that: In S1, the toll station transaction data includes license plate number, vehicle type, entry / exit identification, passage time, and payment type; the gantry transaction data includes license plate number, vehicle type, gantry unique identifier, passage time, and medium type; and the weather data includes temperature, relative humidity, average wind speed, and cumulative rainfall.

3. The data-driven route selection and dynamic feedback method for predicting highway traffic volume according to claim 1, characterized in that: In S1, the feature engineering process specifically involves encoding the fused data. Individual features include the frequency of high-speed travel of vehicles, vehicle type, local administrative region, and energy type. Traffic flow operation features include the difference in travel time, mileage, toll fees, and number of times the alternative route and the shortest route are switched between highways and ordinary roads. Weather features include the average temperature, average relative humidity, average wind speed, and cumulative rainfall for the corresponding time interval and road segment. After one-hot encoding or numerical processing, a 10-15 dimensional feature vector is formed.

4. The data-driven route selection and dynamic feedback method for predicting highway traffic volume according to claim 1, characterized in that: In S3, during the iterative allocation process, a road network impedance update is performed every 100-200 OD pairs allocated. If the road segment load change rate is less than 5% in two adjacent iterations, the iterative allocation of that OD pair is terminated in advance.

5. The data-driven route selection and dynamic feedback method for predicting highway traffic volume according to claim 1, characterized in that: In S4, the conversion coefficients for standard vehicles are as follows: 1.0 for Type I passenger vehicles and Type I freight vehicles; 1.2 for Type II passenger vehicles and Type II freight vehicles; 1.5 for Type III passenger vehicles and Type III freight vehicles; 2.0 for Type IV passenger vehicles and Type IV freight vehicles; and 2.5-4.0 for Type V and above freight vehicles and special-purpose vehicles.

6. The data-driven route selection and dynamic feedback method for predicting highway traffic volume according to claim 1, characterized in that: In S4, the impedance feature vector includes six sub-features: flow characteristics include road segment-time slice load, upstream road segment flow, and downstream road segment flow; saturation characteristics include the ratio of road segment load to dynamic traffic capacity; traffic composition characteristics include the flow proportion of each vehicle type; spatiotemporal characteristics include time interval identifier, road segment location identifier, and weekday attribute; dynamic event characteristics include whether there is a traffic event identifier and event level; and historical state characteristics include the average road segment load of the previous three adjacent time slices.

7. The data-driven route selection and dynamic feedback method for predicting highway traffic volume according to claim 1, characterized in that: In S4, the training samples for the pre-trained LightGBM machine learning model are historical road segment-time slice load data, corresponding impedance feature vectors, and actual travel time data from the past 3-6 months.

8. The data-driven route selection and dynamic feedback method for predicting highway traffic volume according to claim 1, characterized in that: It also includes a model update step: every 7-15 days, the LightGBM path selection model and the LightGBM regression impedance model are incrementally trained using the latest collected multi-source data to update the model parameters to ensure prediction accuracy.