Road network operation situation research method based on multi-source data fusion

Through the road network operation situation research method based on multi-source data fusion, the problem that the existing technology cannot fully reflect the road network operation status is solved, efficient data fusion and analysis are achieved, and the scientificity and accuracy of traffic management are improved.

CN120126306APending Publication Date: 2025-06-10CHINA DESIGN GROUP CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510124728.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art cannot fully and accurately reflect the operating status of the road network, and has not fully explored and applied the large amount of data assets accumulated by highways during the operation of the road network.

Method used

Through a road network operating situation research method based on multi-source data fusion, including selecting road section areas for traffic data and vehicle data acquisition, classifying data according to spatial and temporal dimensions, filtering, interpolation and feature selection, and data fusion using improved Kalman filtering.

Benefits of technology

It improves the accuracy of data mapping, enhances the accuracy of the operation of the transportation road network, provides scientific and accurate decision-making basis, improves the efficiency of road network traffic and ensures public travel safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126306A_ABST
    Figure CN120126306A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent traffic, in particular to a road network operation situation research method based on multi-source data fusion, and the method comprises the following steps: S1, selecting a road section area, and carrying out the collection of traffic data and vehicle data; s2, classifying data collected by a vehicle track layer and a road traffic layer according to two dimensions of space and time to obtain a vehicle GPS track and traffic data, and updating a data set containing a traffic state; s3, filtering and interpolating the updated data, performing feature selection, and performing data balance distribution by using an algorithm; s4, improving the Kalman wave to enable the Kalman wave to read the nearest single sensor; and S5, fusing and outputting the traffic state data. According to the method, data grouping is expanded from one dimension to two dimensions, a vehicle track layer and a road traffic layer are combined, the data mapping accuracy is improved, the Kalman filtering model SCAATF is improved, traffic state data fusion is achieved, and the operation situation of a traffic road network is accurately reflected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation, and particularly to a research method for road network operation status based on multi-source data fusion. Background Art

[0002] At present, the means of highway traffic information collection in service are limited to traditional methods, resulting in a single data collection method, insufficient exploration of data value, and failure to fully utilize the advantages of data such as accessibility, real-time nature, and authenticity. Specific problems are manifested in: on the one hand, the existing data collection system cannot comprehensively and accurately reflect the road network operation status; on the other hand, due to the lack of effective data analysis means, a large amount of data assets accumulated in the highway during road network operation have not been fully explored and applied. Therefore, the key problem currently faced is how to achieve the joint aggregation of highway multi-source monitoring data and conduct in-depth analysis of the road network operation status to solve the existing problems. In this way, a new pattern of highway operation management with digital resources as the core element is constructed, providing scientific and accurate decision-making basis for traffic managers, thereby effectively improving the road network traffic efficiency and ensuring the safety of public travel. Summary of the Invention

[0003] The technical problem to be solved by the present invention is: to achieve the joint aggregation of highway multi-source monitoring data and conduct in-depth analysis of the road network operation status.

[0004] The present application provides a research method for road network operation status based on multi-source data fusion to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solution: A research method for road network operation status based on multi-source data fusion, including the following steps:

[0006] S1. Select a section area to collect traffic data and vehicle data;

[0007] S2. Classify the data collected from the vehicle trajectory layer and the road traffic layer according to two dimensions of space and time to obtain vehicle GPS trajectories and traffic data, and update the data set including traffic status;

[0008] S3. Filter and interpolate the updated data, perform feature selection through principal component analysis PCA, recursive feature elimination RFE, and feature importance FI models, and use an algorithm to balance the data distribution;

[0009] S4. Improve the Kalman wave to make it read the nearest single sensor;

[0010] S5. Fuse traffic status data and output.

[0011] The classification of the spatial dimension in step S2 is as follows: Identify the route mapped by GPS coordinates through the map matching process, and then perform time classification after spatial classification.

[0012] The spatial classification includes the following steps:

[0013] Add timestamps to each section in the selected road network area and modify the data to a tracking-based format;

[0014] Convert the section coordinates and tracking data to GPX format, and the data structure is ['id', 'longitude', 'latitude', 'timestamp'];

[0015] Start from the road network obtained from OSM and model it as a directed graph , where V is the set of vertices (coordinates of a given section), and E is the set of edges (connections between coordinates). Let be the set of coordinate points and timestamps . Let be the current traffic flow, be the road network traffic flow, where . Calculate the distance using the Euclidean distance between P and the directed edge AB, and the distance is defined as follows: .

[0016]

[0017] where P' is the projection of on the section AB, is the Euclidean distance. According to this definition, the distance is equal to the distance in the opposite segment direction . Then, introduce a vertical displacement λ to the road section to reflect the distance between the middle of the road and the middle of the lane. After calculating the distance between P and the road section, measure the score of a path to estimate the error of the algorithm. Based on this method, perform map matching on and to accurately identify the ID of each section for a given P, and then fuse the two sets of data into the same section identification.

[0018] The specific operation of the interpolation in step S3 is as follows: When using sensor data to monitor and control entities (especially vehicles), the problem of data loss caused by sensor failures needs to be considered. Use simple linear interpolation. For each vehicle C = (T, F), where T is the set of trips and F is the set of features; f is a single feature, where f ∈ F and i is its index. The interpolation equation is:

[0019] .

[0020] In step S4, an improved Kalman filter is used to achieve the fusion of traffic state data, which is called SCAATF (Single-Constraint-At-A-Time Filter). The SCAATF filter observes a single latest measurement from a single sensor, and the basic Kalman filter equations and different parameters adopt constant scalar values.

[0021] The core steps of the SCAATF algorithm include initialization, prediction update, and measurement update. First, initialize the state estimate and error covariance, and set the initial speed estimate value and error covariance matrix. In the time update stage, predict the state and error covariance at the current moment according to the state transition matrix. Usually, it is assumed that the state remains unchanged (the state transition matrix A = 1), and a small process noise covariance Q is added. In the measurement update stage, obtain the most recent measurement value from the available sensors, calculate the Kalman gain, correct the state estimate using the residual between the measurement value and the predicted value, and update the error covariance matrix.

[0022] A road network operation situation research system based on multi-source data fusion includes:

[0023] An input module for obtaining traffic information and vehicle information of the selected section and using them as the input data of the proposed system;

[0024] A data spatio-temporal grouping module that classifies the data collected from the vehicle trajectory layer and the road traffic layer according to two dimensions of space and time to obtain vehicle GPS trajectories and traffic data, and updates the dataset containing traffic states;

[0025] A data balancing module that filters and interpolates the updated data, selects features through the principal component analysis PCA, recursive feature elimination RFE, and feature importance FI models, and uses an algorithm for data balanced distribution;

[0026] An SCAATF-based data fusion module that fuses the classified feature data;

[0027] An output module that outputs the fused traffic state data.

[0028] Compared with the prior art, the beneficial effects of the present invention are:

[0029] (1) The present invention expands the data grouping from one dimension (space or time) to two dimensions (space-time), combines the vehicle trajectory layer and the road traffic layer according to two dimensions of time and space. The system input is vehicle GPS trajectories and traffic data, and the result is an updated dataset containing traffic states, which is beneficial to improving the data mapping accuracy;

[0030] (2) The present invention uses three methods, namely PCA, RFE and FI, to select effective data features, and combines SMOTE and KNN algorithms to process unbalanced data. Then, the Tomek link algorithm is applied to find the classes of opposite instances to classify and extract features of the data. The Kalman filter model SCAATF is then improved to realize the fusion of traffic status data and improve the accuracy of reflecting the operation status of the traffic network. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 This is a flow chart of the research method of road network status operation based on multi-source data fusion.

[0032] Figure 2 The attached figure is a flow chart of the SCAATF algorithm. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0034] Example 1

[0035] The present invention provides a road network operation status research method based on multi-source data fusion, which is used to accurately and timely predict and evaluate the upcoming traffic congestion status of important sections of ordinary national and provincial roads. The present invention expands data fusion from one dimension (space or time) to two dimensions (space-time). In the time dimension, traffic network data is analyzed; in the spatial dimension, vehicle trajectory data analysis is implemented to make the data source for road network operation status research more accurate. In order to achieve this goal, firstly, representative road section areas are selected as research objects. These sections should cover different traffic characteristics, such as highways, urban main roads, congestion-prone sections, etc., to ensure the diversity and representativeness of the data. It is specifically achieved through the following steps:

[0036] S1. Collect traffic data and vehicle data for the selected road section area. Including traffic data: including traffic flow, vehicle speed, road occupancy, accident rate, etc., collected through fixed sensors (such as geomagnetic sensors, video surveillance equipment) and mobile sensors (such as vehicle-mounted GPS equipment).

[0037] Vehicle data: including vehicle GPS trajectory data, recording vehicle driving path, speed changes, dwell time and other information.

[0038] S2. Merge the vehicle trajectory layer and the road traffic layer according to two dimensions, namely space and time. The inputs are vehicle GPS trajectories and traffic data. The result of this process is an updated dataset containing traffic states.

[0039] Specifically, the above-mentioned S2 is as follows:

[0040] S21. Spatial grouping; identify the route mapped by GPS coordinates through the map matching process, which specifically includes the following steps:

[0041] S211. Add timestamps to each road segment in the selected road network area and modify the data to a tracking-based format.

[0042] S212. Convert the road segment coordinates and tracking data into GPX format, and the data structure is ['id', 'longitude', 'latitude', 'timestamp'].

[0043] Start from the road network obtained from OSM and model it as a directed graph G(V, E), where V is the set of vertices (coordinates of given road segments) and E is the set of edges (connections between coordinates). Let P_i be the set of coordinate points (x_i, y_i) and timestamps t_i (i = 1, 2,..., n), T_c be the current traffic flow, and T_f be the road network traffic flow, where P_i ∈ T_c and P_i ∈ T_f. Calculate the distance using the Euclidean distance between P and the directed edge AB. The distance is defined as follows:

[0044]

[0045] where P is the projection of P on the road segment AB, and d_e is the Euclidean distance. According to this definition, the distance d_(p,AB) is equal to the distance of the opposite line segment direction d_(p,BA). Then, introduce a vertical displacement λ to the road segment to reflect the distance between the middle of the road and the middle of the lane. After calculating the distance between P and the road segment, measure the score of a path to estimate the error of the algorithm. Based on this method, perform map matching on T_c and T_f to accurately identify the ID of each road segment for a given P. Then, fuse the two sets of data into the same road segment identification to achieve the spatial fusion of vehicle trajectories and road traffic data.

[0046] S22. Temporal grouping. After performing spatial grouping using the map matching method, perform temporal grouping to distinguish vehicle behavior and traffic environment, which specifically includes the following steps:

[0047] S221. Select each element of the vehicle trajectory, including the time granularity, and perform time verification with the traffic data. The time data granularity can be coarse-grained (daily traffic summary) or fine-grained (hourly traffic summary).

[0048] S222. Perform time grouping based on coarse-grained traffic data, and fuse traffic information such as FF (Free flow speed), JF (Jam Factor), SP (speed capped by speed), and SU (speed not capped by speed) into the current road traffic data.

[0049] S3. Filter and interpolate the collected data, perform feature selection through principal component analysis (PCA), recursive feature elimination (RFE), and feature importance (FI) models, and use SMOTE and KNN algorithms to balance the data distribution. The specific steps are as follows:

[0050] S31. Filtering. The premise considered during analysis is that even if only vehicle sensor data is used, it is possible to provide valuable information about traffic behavior. Under this premise, all variables with problems such as outliers, conflicts, incompleteness, ambiguity, correlation, differences, etc., or those that cannot reflect traffic behavior are removed from the collected data.

[0051] S32. Interpolation. When using sensor data to monitor and control entities (especially vehicles), the problem of data loss caused by sensor failures needs to be considered. Simple linear interpolation is used. For each vehicle C = (T, F), where T is the set of trips and F is the set of features; f is a single feature, where f ∈ F and i is its index. The interpolation equation is:

[0052]

[0053] S33. Feature selection. Select effective data features through various machine learning methods such as principal component analysis (PCA), recursive feature elimination (RFE), and feature importance (FI), reduce the data dimension, improve the model training efficiency, and perform feature set identification.

[0054] S331. Principal component analysis (PCA), which is used to extract a set of relevant features. This process identifies the most variable information from a multivariate dataset and represents it as a set of new features - principal components (PCs). The PCs represent the directions of the greatest data variation.

[0055] S332. Recursive feature elimination (RFE) is used to select features suitable for the model to obtain higher accuracy. RFE ranks these features according to the feature importance attribute of the model and recursively eliminates possible dependencies and collinearity in the model.

[0056] S333. Feature importance (FI). This classifier calculates the relative importance of each feature. This technique calculates the probability of reaching a node as the number of samples reaching the node divided by the total number of samples. The higher the value, the more important the feature.

[0057] S34. Balance the data. To address the problem of class imbalance in the dataset, the following methods are adopted:

[0058] SMOTE algorithm: By synthesizing minority class samples, increase the number of minority classes in the dataset and improve the data distribution.

[0059] KNN algorithm: Combine the K-Nearest Neighbors algorithm to find similar observations of the minority class and create synthetic samples in the space.

[0060] Tomek Links algorithm: Find and remove pairs of opposite instances of classes to further optimize the data distribution and improve the generalization ability of the model.

[0061] Combine two techniques to handle imbalanced data. SMOTE (The Synthetic Minority Over-sampling Technique) uses the k-Nearest Neighbors (KNN) algorithm to find similar observations of the minority class and randomly selects a KNN to create synthetic samples in the space. Next, apply the Tomek Links algorithm to find pairs of opposite instances of classes.

[0062] S4. Improve the Kalman wave to make it read the nearest single sensor; called SCAATF (Single-Constraint-At-A-Time Filter). This method uses a single latest measurement from any available sensor (updating its estimate according to the characteristics of the observed sensor, i.e., the variance of the measurement) and the state estimate from the previous step.

[0063] The specific content of S4 is as follows: To achieve the fusion of traffic state data, the present invention proposes an improved Kalman filter SCAATF (Single-Constraint-At-A-Time Filter). The core idea of this filter is to observe the latest measurement from a single sensor and combine the state estimate from the previous step to update the traffic state information in real time. Its basic steps are as follows:

[0064] State prediction: Predict the state at the current moment based on the state estimate at the previous moment and the system dynamic model.

[0065] Measurement update: Combine the sensor measurement value at the current moment to correct the predicted state and obtain a more accurate state estimate.

[0066] Parameter optimization: Optimize the performance of the filter by adjusting parameters such as the Kalman gain to ensure its sensitivity and adaptability to traffic state changes.

[0067] The SCAATF filter reads the nearest single sensor in the measurement update step instead of waiting for all sensors to obtain a complete dataset.

[0068] The SCAATF filter can effectively handle the fusion problem of multi-source data by introducing a single constraint condition, reduce data noise and uncertainty, and improve the accuracy of traffic state estimation.

[0069] The SCAATF filter observes a single latest measurement from a single sensor. The basic Kalman filter equations and different parameters adopt constant scalar values as follows:

[0070] Current state variable: ;

[0071] State covariance: ;

[0072] Kalman filter gain: ;

[0073] Use the observation vector to perform state update: ;

[0074] Error covariance update: .

[0075] A = 1: There is no reason to believe that the traffic condition has changed unless there is a new measurement method.

[0076] B = 0: There are no known external control input factors affecting the state measurement.

[0077] P value = sensor-specific error variance: Each sensor has its unique specific error variance value.

[0078] Q = 1e-5: There is no known process noise.

[0079] S5. Integrate traffic state data and output.

[0080] S5.1 Scheduling strategy formulation

[0081] According to the traffic state prediction results, formulate a road network scheduling strategy to optimize traffic flow allocation and relieve congestion:

[0082] Construction of the scheduling parameter library: Based on historical scheduling experience, establish a parameter library containing various scheduling strategies, covering aspects such as traffic flow control, signal light optimization, and path planning.

[0083] Matching of scheduling constraint parameters: Use the difference between the predicted traffic state and the scheduling demand parameters as a constraint condition to match the optimal combination of scheduling parameters from the parameter library.

[0084] S5.2 Global optimization

[0085] Within the optimized search space of the matching scheduling parameters, comparison is carried out through global optimization algorithms (such as genetic algorithms and particle swarm optimization) to determine the optimal road network scheduling parameters, maximize the traffic operation efficiency, and output data.

[0086] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

Claims

1. A road network operation status research method based on multi-source data fusion, characterized in that: The following steps are involved: S1. Select a road section area to collect traffic data and vehicle data; S2, classify the data collected from the vehicle trajectory layer and the road traffic layer according to the two dimensions of space and time, obtain the vehicle GPS trajectory and traffic data, and update the data set containing the traffic status; S3, filter and interpolate the updated data, perform feature selection through principal component analysis (PCA), recursive feature elimination (RFE) and feature importance (FI) model, and use algorithms to balance data distribution; S4, improve the Kalman wave to read the nearest single sensor; S5. Integrate traffic status data and output.

2. The method for studying road network operation status based on multi-source data fusion according to claim 1 is characterized in that: The classification of the spatial dimension in step S2 is as follows: identifying the route mapped by the GPS coordinates through a map matching process, and performing time classification after spatial classification.

3. The method for studying road network operation status based on multi-source data fusion according to claim 2 is characterized in that: The spatial classification comprises the following steps: Add timestamps to each road segment in the selected road network area and modify the data to a trace-based format; Convert the road segment coordinates and tracking data to GPX format, the data structure is ['id', 'longitude', 'latitude', 'timestamp']; Start with a road network obtained from OSM and model it as a directed graph , where V is the set of vertices (the coordinates of a given road segment) and E is the set of edges (the connections between coordinates). is the coordinate point and timestamp A collection of is the current traffic flow, is the road network traffic flow, where , The distance is calculated using the Euclidean distance from P to the directed edge AB, and the distance is defined as follows: Among them, P' is The projection on the road section AB, is the Euclidean distance.

4. The method for studying road network operation status based on multi-source data fusion according to claim 1 is characterized in that: The specific interpolation operation in step S3 is: when using sensor data to monitor and control entities (especially vehicles), it is necessary to consider the problem of missing data caused by sensor failure. Using simple linear interpolation, for each vehicle C=(T, F), where T is the trip set, F is the feature set; f is a single feature, the interpolation equation is: Where f∈F, i is its index.

5. The method for studying road network operation status based on multi-source data fusion according to claim 1 is characterized in that: In step S4, an improved Kalman filter is used to realize the fusion of traffic status data, which is called SCAATF (Single-Constraint-At-A-Time Filter). The SCAATF filter observes the single latest measurement value from a single sensor, and the basic Kalman filter equation and different parameters use constant scalar values.

6. The method for studying road network operation status based on multi-source data fusion according to claim 1 is characterized in that: The core steps of the SCAATF algorithm include initialization, prediction update and measurement update. The initialization is specifically as follows: initializing state estimation and error covariance, setting initial velocity estimation value and error covariance matrix. In the time update phase, the state and error covariance at the current moment are predicted according to the state transfer matrix. It is usually assumed that the state remains unchanged and a small process noise covariance Q is added. In the measurement update phase, the most recent measurement value is obtained from the available sensor, the Kalman gain is calculated, the state estimation is corrected using the residual between the measurement value and the prediction value, and the error covariance matrix is ​​updated.

7. A road network operation status research system based on multi-source data fusion, characterized in that: include: An input module, used to obtain the traffic information and vehicle information of the selected road section and use them as input data of the proposed system; The data spatiotemporal grouping module classifies the data collected from the vehicle trajectory layer and the road traffic layer according to the two dimensions of space and time, obtains the vehicle GPS trajectory and traffic data, and updates the data set containing the traffic status; The data balancing module filters and interpolates the updated data, performs feature selection through principal component analysis (PCA), recursive feature elimination (RFE) and feature importance (FI) model, and uses algorithms to balance data distribution; The data fusion module based on SCAATF is used to fuse the classified feature data; Output module, outputs the fused traffic status data.

Citation Information

Cited By

  • Intelligent agent and urban road interaction real-time scene analysis method

    CN120340262A

  • A real-time scene analysis method for the interaction between intelligent agents and urban roads

    CN120340262B