Traffic flow state identification method based on multi-source heterogeneous data source

Through the traffic flow state recognition method of multi-source heterogeneous data sources, combined with the spatiotemporal graph convolutional network and ensemble learning model, the problem of incomplete data acquisition in traditional methods is solved, and efficient and accurate recognition and real-time management of highway traffic flow status are achieved.

CN120726802AActive Publication Date: 2025-09-30广西计算中心有限责任公司

Patent Information

Application Number
CN202510822907.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-30
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional traffic flow status recognition methods rely on fixed vehicle inspection data or a single data source. They suffer from incomplete data acquisition, poor real-time performance, and weak anti-interference capabilities, making it difficult to meet the needs of modern highway emergency command systems.

Method used

A traffic flow state recognition method based on multi-source heterogeneous data sources is adopted. By collecting and classifying traffic information data, pre-processing it and inputting it into the classification model, the road network model is modeled in combination with the spatiotemporal graph convolutional network model, and the final traffic state is comprehensively corrected. The classification results of the SVM and decision tree models are integrated through an ensemble learning method, and weights are assigned to calculate the final traffic state.

Benefits of technology

It improves the accuracy and real-time performance of traffic flow status identification, enhances anti-interference ability, optimizes traffic management strategies, and improves resource allocation efficiency and management level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726802A_ABST
    Figure CN120726802A_ABST
Patent Text Reader

Abstract

The invention provides a traffic flow state identification method based on a multi-source heterogeneous data source. The method comprises the following steps: collecting and classifying traffic information data of road traffic; preprocessing the traffic information data; respectively inputting the preprocessed traffic information data into corresponding classification models according to types; outputting an intermediate traffic state corresponding to each type; weight distribution is carried out on each type of intermediate traffic state; and calculating to obtain a final traffic state. The method has the advantages that a customized feature extraction and conversion method is adopted according to the characteristics of different types of sensor data, semantic information of different types of data can be better reserved, and the quality and effect of data fusion are improved; the performance of different models on the verification set is integrated, and the voting weight is dynamically adjusted according to the accuracy of the models, so that the final prediction result is more accurate and reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of road traffic artificial intelligence, and in particular relates to a traffic flow state recognition method based on multi-source heterogeneous data sources. Background Art

[0002] With the acceleration of urbanization and the dramatic increase in vehicle ownership, highways, as vital transportation arteries, are experiencing increasing traffic volumes. Traffic congestion and accidents are becoming frequent, posing severe challenges to the safe and efficient operation of highways. Traditional traffic flow state recognition methods rely heavily on fixed vehicle inspection data or a single data source. These methods suffer from incomplete data acquisition, poor real-time performance, and weak anti-interference capabilities, making them inadequate for meeting the requirements of modern highway emergency command systems. Therefore, developing an efficient, accurate, and multi-source data fusion method for highway traffic flow state recognition is crucial. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention provides a traffic flow state recognition method based on multi-source heterogeneous data sources, which is particularly suitable for processing traffic flow state recognition method for multi-source heterogeneous data sources.

[0004] The technical solution adopted by the present invention is as follows: First, a method for identifying traffic flow status based on multi-source heterogeneous data sources is provided, comprising the following steps:

[0005] Collect and classify traffic information data of road traffic;

[0006] Preprocessing the traffic information data;

[0007] Inputting the pre-processed traffic information data into corresponding classification models according to their types;

[0008] Output the intermediate traffic status corresponding to each type;

[0009] assigning a weight to each type of the intermediate traffic state;

[0010] Calculate the final traffic status.

[0011] Furthermore, the method further comprises the following steps:

[0012] Modeling road networks through spatiotemporal graph convolutional network models;

[0013] Comprehensively correcting the final traffic state of the target road in combination with the real-time traffic states of other roads in the road network;

[0014] The final traffic status is uploaded to the traffic management system and / or the user terminal.

[0015] Furthermore, the types of the traffic information data include numerical values ​​and images; the traffic information data include average vehicle speed, traffic volume and lane occupancy rate.

[0016] Furthermore, preprocessing the traffic information data includes the following steps:

[0017] Data cleaning is performed using spatial clustering algorithms;

[0018] Perform data conversion;

[0019] Perform unified representation of attribute domains;

[0020] Central cluster analysis was performed using the K-means algorithm;

[0021] Data dimensionality reduction is performed using principal component analysis algorithm;

[0022] Data compression is performed using differential encoding.

[0023] Furthermore, performing unified attribute domain representation includes the following steps:

[0024] For the numerical data in the traffic information data, directly retain its numerical attributes;

[0025] Extracting features from the image data in the traffic information data and converting the features into a numerical vector form;

[0026] Establish a data dictionary to clarify the name, data type, value range and meaning of each feature field.

[0027] Furthermore, in the central cluster analysis, the equation Calculate the centroid, where is the centroid of the k-th cluster, is the set of data points of the kth cluster, Cluster The number of data points in Cluster A data point in .

[0028] Furthermore, the training of the classification model includes the following steps:

[0029] Collect historical traffic information data;

[0030] Preprocessing the historical traffic information data;

[0031] Labeling the historical traffic information data according to classification rules;

[0032] The classification model includes an SVM model and a decision tree model, the SVM model is constructed based on the average vehicle speed, and the decision tree model is constructed based on the vehicle flow and the lane occupancy rate;

[0033] Inputting the historical traffic information data into the SVM model and the decision tree model for training respectively according to classification;

[0034] The classification results of the SVM model and the decision tree model are integrated through an ensemble learning method.

[0035] Furthermore, the SVM model optimizes the classification boundary by maximizing the objective function, which is ,in is the Lagrange multiplier, is the sample label, is the kernel function, is the sample size.

[0036] Furthermore, the lane occupancy rate Through the equation The average lane occupancy is obtained by Eq. Get, among them is the number of traffic flows in the i-th lane, Y is the number of single-direction traffic flows, and N is the number of lanes.

[0037] Furthermore, the final traffic state Through the equation ,in is the weight of the average speed of the vehicle, is the weight of the traffic flow, is the weight of the lane occupancy rate, is the intermediate traffic state of the average vehicle speed, is the intermediate traffic state of the traffic flow, The intermediate traffic state is the lane occupancy rate.

[0038] The advantages and positive effects of the present invention are as follows: due to the adoption of the above technical solution, customized feature extraction and conversion methods are adopted according to the characteristics of different types of sensor data, which can better retain the semantic information of different types of data and improve the quality and effect of data fusion; the performance of different models on the verification set is integrated, and the voting weight is dynamically adjusted according to the accuracy of the model, so that the final prediction result is more accurate and reliable; the accuracy and real-time performance of traffic flow state recognition are improved, and fast and accurate data support is provided for the highway emergency command system; the anti-interference ability is enhanced, and the error and interference caused by a single data source are reduced; the formulation and implementation of traffic management strategies are optimized, and the allocation efficiency and management level of traffic resources are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flow chart of a traffic flow state recognition method based on multi-source heterogeneous data sources according to an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The present disclosure is described more fully below with reference to the accompanying drawings, which illustrate exemplary embodiments of the present disclosure. The technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present disclosure.

[0041] like Figure 1 As shown, the present invention provides a traffic flow state identification method based on multi-source heterogeneous data sources, comprising the following steps:

[0042] S100, collecting and classifying traffic information data of road traffic;

[0043] Classification algorithms are used to preliminarily screen potential data sources and select parameters that are highly correlated with traffic flow status, including but not limited to vehicle speed, traffic volume, lane occupancy, etc.

[0044] S200, pre-processing traffic information data;

[0045] S300, inputting the pre-processed traffic information data into corresponding classification models according to their types;

[0046] S400: Outputting intermediate traffic states corresponding to various types of traffic, including smooth traffic, slow traffic, and congested traffic.

[0047] S500, assigning weights to various types of intermediate traffic states;

[0048] Different weights are assigned to the three classification results of vehicle speed, traffic volume, and lane occupancy. The classification weights are assigned based on historical data, expert experience, and specific road conditions. The sum of all weights is 1.

[0049] S600: Calculate and obtain the final traffic status.

[0050] By using the above method, the optimal processing path is adapted for different data types, retaining the key information of the original data; through independent processing through classification, the anomaly of a single data source is avoided from contaminating the global results; and each data stream is processed in parallel to effectively reduce system latency.

[0051] In order to solve the problem that traditional methods ignore the traffic flow conduction effect between road sections and cannot break through the spatial limitations of single-point state recognition, an implementation method is provided in this embodiment.

[0052] In one embodiment, the following steps are further included:

[0053] Modeling road networks through spatiotemporal graph convolutional network models;

[0054] The final traffic status of the target road is comprehensively corrected by combining the real-time traffic status of other roads in the road network;

[0055] The final traffic status is uploaded to the traffic management system and / or user terminal.

[0056] The recognition results are fed back to the traffic management system or user terminal in real time, and the time series analysis algorithm is used to predict the trend of traffic status changes in the future, providing a scientific basis for traffic scheduling, route planning, information release, etc.

[0057] Using this method, the ST-GCN is used to capture the spatiotemporal propagation patterns of traffic flow and predict the congestion diffusion trend. The results of the current section are corrected by integrating the status of adjacent sections, effectively reducing the false alarm rate. This provides a certain warning window period for traffic management systems and end users.

[0058] In one embodiment, the types of traffic information data include numerical values ​​and images; the traffic information data include average vehicle speed, traffic volume, and lane occupancy rate.

[0059] In order to solve the problem of incompatible scales of multi-source data, an implementation method is provided in this embodiment.

[0060] In one embodiment, preprocessing the traffic information data includes the following steps:

[0061] Data cleaning is performed using spatial clustering algorithms;

[0062] Use density-based noise spatial clustering algorithm to identify and remove noise points and outliers in the data set. Use spatial clustering algorithm to identify noise points and set the neighborhood radius and the minimum number of points MinPts, data points is a noise point if and only if ,in for of - A collection of points within a domain.

[0063] Perform data conversion;

[0064] Use data transformation techniques including but not limited to linear transformation and normalization to ensure that data from different data sources are compared and clustered on the same scale.

[0065] Perform unified characterization of the attribute domain;

[0066] Unify the data collected from different sensors (such as fixed vehicle detectors, GPS devices, video cameras) in the attribute domain, including but not limited to unifying all speed data to kilometers per hour (km / h), all traffic flow data to vehicles per hour (veh / h), and lane occupancy rate data to percentage (% ).

[0067] Perform central clustering analysis through the K-means algorithm;

[0068] Perform data dimensionality reduction through the principal component analysis algorithm;

[0069] Assume that the dimension of the original data is d, and the dimension of the data after dimensionality reduction is d′ (d′ < d). Then the goal of PCA is to minimize the reconstruction error. where V is the transformation matrix, is the projection matrix projected onto the principal component space.

[0070] Perform data compression through differential coding.

[0071] The preprocessing methods also include spatio-temporal synchronization of multi-source data. For the GPS data delay problem, a timestamp-based compensation method is adopted. Record the timestamp of each GPS data point, compare it with the current system time, calculate the delay time, and perform interpolation processing on the GPS data according to the delay time to estimate the GPS data position at the correct time point. Adopt the linear interpolation method to calculate the position coordinates at the correct time point based on two known GPS data points before and after the delay time.

[0072] The preprocessing methods also include the video sampling rate and vehicle detector data alignment strategy. For the problem of misalignment between the video sampling rate and vehicle detector data, adopt the timestamp alignment and resampling methods. Add timestamps to the video frames and vehicle detector data to ensure that each data point has accurate time information. Align the video frames and vehicle detector data according to the timestamps. For the case of inconsistent sampling rates, adopt the resampling method to resample the high-sampling-rate data to the same sampling rate as the low-sampling-rate data, or adopt the interpolation method to estimate the values of missing data points based on the values of adjacent data points to ensure the consistency of the data in the time dimension.

[0073] Adopt the above methods to establish a comparable data benchmark through unified characterization of the attribute domain; differential coding compression reduces the transmission bandwidth.

[0074] To solve the problems of feature explosion caused by directly inputting original pixels in traditional methods and the semantic gap between video data and numerical vectors, an implementation method is provided in this embodiment.

[0075] In one embodiment, performing unified attribute domain representation includes the following steps:

[0076] For numerical data in traffic information data, directly retain its numerical attributes;

[0077] Extract features from image data in traffic information data and convert them into numerical vector form;

[0078] Establish a data dictionary to clarify the name, data type, value range and meaning of each feature field to ensure data consistency and comprehensibility.

[0079] Using this method, the vehicle's motion trajectory is analyzed from video frames to capture micro-behaviors such as sudden acceleration or lane changes; the influence of brightness fluctuations caused by weather is eliminated through feature vector conversion.

[0080] In order to solve the problem that computing efficiency cannot be optimized under massive data and resources are wasted due to full processing of all sensor data, an implementation method is provided in this embodiment.

[0081] In one embodiment, the center cluster is selected (k=1,2,...,K) as representative data sources, where X={x1,x2,...,x n} is the original data set, n is the number of data points, and the central cluster analysis is performed by equation Calculate the centroid, where is the centroid of the k-th cluster, is the set of data points of the kth cluster, Cluster The number of data points in Cluster A data point in .

[0082] The above method is used to screen the centroid representative points to filter out redundant data and reduce calculation time; the K-means algorithm can more accurately select representative data sources based on the intrinsic characteristics of the data when clustering.

[0083] In order to solve the problems of being unable to reconcile the contradiction between rule-driven and data-driven, the poor interpretability of pure machine learning models, and the insufficient flexibility of pure rule models, an implementation method is provided in this embodiment.

[0084] In one embodiment, training the classification model includes the following steps:

[0085] Collect historical traffic information data;

[0086] Preprocess historical traffic information data;

[0087] Label historical traffic information data according to classification rules;

[0088] Based on vehicle speed, traffic volume, and lane occupancy, corresponding thresholds are set and classification rules are constructed. A smooth traffic state is determined when the average vehicle speed is above a certain threshold (e.g., 80 km / h), the lane occupancy is low (e.g., less than 20%), and the traffic volume is moderate (e.g., less than 60% of the road's design capacity). A congested traffic state is determined when the vehicle speed is below a certain threshold (e.g., 40 km / h), the traffic volume exceeds a certain percentage of the road's design capacity (e.g., 80%), and the lane occupancy is above a certain threshold (e.g., 40%).

[0089] The classification model includes the SVM model and the decision tree model. The SVM model is constructed based on the average vehicle speed, and the decision tree model is constructed based on the traffic flow and lane occupancy rate.

[0090] The decision tree model uses the C4.5 decision tree algorithm to construct a binary tree structure, selects the best splitting attribute according to the information gain ratio, and makes independent judgments on each indicator.

[0091] The historical traffic information data is input into the SVM model and decision tree model for training according to the classification;

[0092] The classification results of the SVM model and the decision tree model are integrated through the ensemble learning method.

[0093] The SVM model and decision tree model are used in parallel, and the results are integrated using a voting method. The training dataset is fed into both the SVM model and the decision tree model for training. During the prediction phase, each new input sample is predicted using both the trained SVM model and the decision tree model, yielding predictions from both models. A vote is then taken based on the predictions of the two models. If the predictions of the two models agree, that result is used as the final prediction. If the predictions of the two models disagree, a weighted vote is performed based on the accuracy of the two models on the validation set, with the predictions of the model with the higher accuracy being given a higher weight, resulting in the final prediction.

[0094] Using the above method, the decision tree explicitly embeds traffic domain knowledge, and the SVM implicitly learns complex speed patterns; the voting mechanism makes the overall accuracy exceed that of any single model.

[0095] In order to solve the problem that the traditional threshold method cannot identify the linear inseparability of the "high-speed but congested" scenario and the speed mutation state, an implementation method is provided in this embodiment.

[0096] In one embodiment, the SVM model optimizes the classification boundary by maximizing the objective function, which is ,in is the Lagrange multiplier, is the sample label, is the kernel function, is the sample size.

[0097] For the classification of vehicle speed, the objective function is expressed as ,in is the Lagrange multiplier, For samples Label, is the kernel function, is the sample size, used to calculate the sample and The similarity between them, commonly used kernel functions include linear kernel, polynomial kernel, radial basis function (RBF) kernel, etc.

[0098] Using this method, the hypersurface classification boundary is constructed through the RBF kernel function, which improves the recognition rate of sudden deceleration; grid search ensures that the model is continuously optimized as road conditions change.

[0099] In order to solve the problem that traditional methods ignore lane differences and distort the average occupancy rate of the entire road section, an implementation method is provided in this embodiment.

[0100] In one embodiment, lane occupancy Through the equation The average lane occupancy is obtained by Eq. Get, among them is the number of traffic flows in the i-th lane, Y is the number of single-direction traffic flows, and N is the number of lanes.

[0101] An object detection algorithm (such as the YOLO series) is used to detect vehicles in the video and obtain vehicle location information (such as bounding box coordinates). Then, based on the pre-defined lane line positions, the system determines whether the detected vehicle is within the lane and counts the number of vehicles in the lane. Furthermore, the system uses the video frame rate to calculate the number of vehicles passing through a lane per unit time, thereby determining the lane occupancy rate. Lane occupancy data primarily comes from video analysis, which uses video data to perform data clearing and lane identification.

[0102] By adopting the above method, lane-by-lane calculation can prevent lanes with dense truck traffic from obscuring overall smooth traffic; the lanes at the source of congestion can be identified through the average lane occupancy rate.

[0103] In order to solve the problem of insufficient accuracy of single-dimensional traffic status, an implementation method is provided in this embodiment.

[0104] In one embodiment, the final traffic state Through the equation ,in is the weight of the average vehicle speed, is the weight of traffic flow, is the weight of lane occupancy, is the intermediate traffic state with the average vehicle speed, is the intermediate traffic state of traffic flow, is the intermediate traffic state of lane occupancy.

[0105] Using the above method, the weighted results are more interpretable than the black box model output.

[0106] The following describes the contents involved in the above embodiment in conjunction with a preferred embodiment.

[0107] On a certain highway section, a data set containing fixed vehicle detector data, mobile detection data (GPS data), and video image data was obtained. The fixed vehicle detector data at time T1 is v1=40km / h, q1=120 vehicles / hour, the mobile detection data is GPS data 1, and the video image data o1=30%; the fixed vehicle detector data at time T2 is v2=20km / h, q2=180 vehicles / hour, the mobile detection data is GPS data 2, and the video image data o2=45%; the fixed vehicle detector data at time Tn is v n =10km / h, q n = 250 vehicles / hour, the mobile detection data is GPS data n, video image data o n=60%. A density-based noise spatial clustering algorithm was used to identify and remove noise points and outliers from the dataset. GPS data 2 at time T2 was found to be anomaly and was therefore removed. Principal component analysis was used to reduce the data dimension while retaining key information. Assuming the original data had three dimensions (speed, traffic flow, and lane occupancy), the reduced data still had three dimensions, but the data had been optimized. Thresholds for vehicle speed, traffic flow, and lane occupancy were set to determine traffic flow status. Congestion was determined when the average vehicle speed was below a certain threshold (e.g., 30 km / h), lane occupancy was high (e.g., greater than 50%), and traffic volume was heavy (e.g., exceeding 80% of the road's design capacity). An SVM classification model was used to independently assess vehicle speed, traffic flow, and lane occupancy. The preprocessed data was input into the SVM model to obtain preliminary classification results. The intermediate result obtained at time T1 was that the vehicle speed was above the threshold, the lane occupancy was below the threshold, and traffic flow was moderate, indicating a smooth traffic flow. At time T2 (after removing abnormal data): the vehicle speed is below the threshold, the lane occupancy rate is above the threshold, and the traffic volume is high, resulting in a preliminary judgment of congestion (requiring further integration and judgment). At time Tn: the vehicle speed is far below the threshold, the lane occupancy rate is far above the threshold, and the traffic volume is extremely high, resulting in a preliminary judgment of congestion (requiring further integration and judgment). Based on historical data and expert experience, different weights are assigned to the three classification results of vehicle speed, traffic volume, and lane occupancy (for example, 0.4 for speed, 0.3 for traffic volume, and 0.3 for occupancy). The three classification results and their weights are weighted and summed to obtain the final traffic flow state probability. The weighted results at both times T2 and Tn indicate congestion. The identification results are fed back to the traffic management system in real time, which immediately transmits the information to the emergency command system, triggering appropriate emergency response measures, such as issuing traffic guidance information.

[0108] Based on the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0109] An electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the traffic flow state identification method based on multi-source heterogeneous data sources provided by the present disclosure.

[0110] Electronic device is intended to refer to various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device may also refer to various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended to be examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0111] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the traffic flow state identification method based on multi-source heterogeneous data sources provided by the present disclosure.

[0112] Various embodiments of the present disclosure may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0113] A computer program product includes a computer program / instruction. When the computer program / instruction is executed by a processor, the traffic flow state identification method based on multi-source heterogeneous data sources provided by the present disclosure is implemented.

[0114] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0115] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0116] The embodiments of the present invention are described in detail above, but the contents described are only preferred embodiments of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A traffic flow state recognition method based on multi-source heterogeneous data sources, characterized in that: The following steps are involved: Collect and classify traffic information data of road traffic; Preprocessing the traffic information data; Inputting the pre-processed traffic information data into corresponding classification models according to their types; Output the intermediate traffic status corresponding to each type; assigning a weight to each type of the intermediate traffic state; Calculate the final traffic status.

2. The traffic flow state identification method based on multi-source heterogeneous data sources according to claim 1 is characterized in that: The following steps are also included: Modeling road networks through spatiotemporal graph convolutional network models; Comprehensively correcting the final traffic state of the target road in combination with the real-time traffic states of other roads in the road network; The final traffic status is uploaded to the traffic management system and / or the user terminal.

3. The traffic flow state identification method based on multi-source heterogeneous data sources according to claim 1 is characterized by: The types of the traffic information data include numerical values ​​and images; the traffic information data include average vehicle speed, traffic volume and lane occupancy rate.

4. The traffic flow state identification method based on multi-source heterogeneous data sources according to claim 3 is characterized in that: Preprocessing the traffic information data includes the following steps: Data cleaning is performed using spatial clustering algorithms; Perform data conversion; Perform unified representation of attribute domains; Central cluster analysis was performed using the K-means algorithm; Data dimensionality reduction is performed using principal component analysis algorithm; Data compression is performed using differential encoding.

5. The traffic flow state identification method based on multi-source heterogeneous data sources according to claim 4 is characterized in that: The unified representation of attribute domains includes the following steps: For the numerical data in the traffic information data, directly retain its numerical attributes; Extracting features from the image data in the traffic information data and converting the features into a numerical vector form; Establish a data dictionary to clarify the name, data type, value range and meaning of each feature field.

6. The traffic flow state identification method based on multi-source heterogeneous data sources according to claim 4 is characterized by: In central cluster analysis, the equation Calculate the centroid, where is the centroid of the k-th cluster, is the set of data points of the kth cluster, Cluster The number of data points in Cluster A data point in .

7. The method for identifying traffic flow status based on multi-source heterogeneous data sources according to any one of claims 3 to 6, characterized in that: The training of the classification model includes the following steps: Collect historical traffic information data; Preprocessing the historical traffic information data; Labeling the historical traffic information data according to classification rules; The classification model includes an SVM model and a decision tree model, the SVM model is constructed based on the average vehicle speed, and the decision tree model is constructed based on the vehicle flow and the lane occupancy rate; Inputting the historical traffic information data into the SVM model and the decision tree model for training respectively according to classification; The classification results of the SVM model and the decision tree model are integrated through an ensemble learning method.

8. The method for identifying traffic flow status based on multi-source heterogeneous data sources according to claim 7 is characterized by: The SVM model optimizes the classification boundary by maximizing the objective function, which is ,in is the Lagrange multiplier, is the sample label, is the kernel function, is the sample size.

9. The method for identifying traffic flow status based on multi-source heterogeneous data sources according to claim 3 is characterized by: Lane occupancy rate Through the equation The average lane occupancy is obtained by Eq. Get, among them is the number of traffic flows in the i-th lane, Y is the number of single-direction traffic flows, and N is the number of lanes.

10. The method for identifying traffic flow status based on multi-source heterogeneous data sources according to claim 3, characterized in that: The final traffic state Through the equation ,in is the weight of the average speed of the vehicle, is the weight of the traffic flow, is the weight of the lane occupancy rate, is the intermediate traffic state of the average vehicle speed, is the intermediate traffic state of the traffic flow, The intermediate traffic state is the lane occupancy rate.

Citation Information

Patent Citations

  • Traffic event identification method and device

    CN108091131A

  • Emergency lane dynamic management method and system

    CN119723878A

  • Highway traffic flow state recognition method based on deep neural network

    WO2020220439A1

Cited By

  • Industrial internet-oriented computer data feature extraction method

    CN121071465A