A hierarchical classification and processing system for multi-source traffic data on highways

By establishing a hierarchical and classified processing system for multi-source traffic data on highways, the problems of insufficient dynamism and intelligence in data processing in existing technologies have been solved. This system enables rapid response to emergency events and efficient fusion of multi-source data, thereby improving the accuracy of data processing and the efficiency of resource utilization.

CN121122025BActive Publication Date: 2026-03-13SHANDONG EXPRESSWAY INFORMATION GRP CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-03-13

Smart Images

  • Figure CN121122025B_ABST
    Figure CN121122025B_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent transportation systems and discloses a hierarchical classification and processing system for multi-source traffic data on highways. The system includes: a data acquisition standardization module for collecting and preprocessing multi-source traffic data; a data source quality assessment module for evaluating the credibility of data sources; a scenario urgency calculation module for calculating traffic flow anomaly, vehicle speed anomaly, and meteorological risk scores, and identifying emergency events; a data classification module for calculating the comprehensive priority score of the data and obtaining a hierarchical dataset using a threshold segmentation method; a data classification module for adaptively classifying the data; a data fusion module for identifying the same traffic parameters and performing conflict detection, and fusing conflicting data; and a result output module for obtaining the processing results using a hierarchical classification output and quality feedback mechanism. This invention achieves intelligent and refined processing of multi-source traffic data on highways, improving the response speed to emergency events and the accuracy of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent transportation systems, and more specifically, to a hierarchical classification and processing system for multi-source traffic data on highways. Background Technology

[0002] With the rapid development of my country's expressway network and the continuous improvement of intelligent transportation systems, expressway traffic data exhibits characteristics of being multi-sourced, heterogeneous, massive, real-time, and of varying quality. These data sources include high-definition video surveillance, microwave detectors, geomagnetic sensors, ETC gantry systems, meteorological monitoring stations, mobile communication signaling, and vehicle-mounted GPS devices, among others. How to effectively classify and process this multi-source traffic data to extract valuable traffic information has become a key technical issue for intelligent expressway management.

[0003] Existing traffic data processing technologies suffer from the following shortcomings: Firstly, data grading mechanisms lack dynamism, mostly employing static rule-based classification methods. This fails to dynamically adjust data processing priorities according to changes in actual traffic conditions, leading to delays in emergency response. For example, when a traffic accident occurs on a highway, relevant emergency data may be submerged in massive amounts of routine data, impacting emergency response speed. Secondly, data quality assessment methods are simplistic, typically relying on basic judgments based solely on data integrity or timeliness, neglecting the historical reliability of data sources, current operational status, and cross-validation between data. This results in erroneous data from low-quality data sources interfering with system judgments. Thirdly, classification and processing strategies are crude, often applying a uniform processing flow to all data of the same type, failing to refine processing based on differences in data characteristics, and lacking the ability to automatically identify data patterns and configure optimal processing algorithms based on statistical, temporal, and spatial characteristics.

[0004] The lack of effective conflict resolution mechanisms in multi-source data fusion means that simple voting or weighted averaging is insufficient to obtain accurate results when information from different data sources contradicts each other, failing to adequately handle data conflicts and the uncertainty of quantification results. Existing technologies mostly employ centralized processing architectures, which struggle to simultaneously meet the millisecond-level real-time response requirements for urgent data and the in-depth analysis needs of routine data. Furthermore, existing technologies lack adaptive resource scheduling mechanisms for data with varying priorities and processing complexities, failing to dynamically allocate computing resources and storage space based on data classification results and system resource status, leading to inefficient resource utilization. These technical problems necessitate a comprehensive and intelligent multi-source traffic data classification and processing system for highways to address these issues. Summary of the Invention

[0005] This invention provides a hierarchical classification and processing system for multi-source traffic data on highways, which solves the technical problems of lack of dynamism in data classification mechanisms, single data quality assessment methods, and crude classification and processing strategies in related technologies.

[0006] This invention provides a hierarchical classification and processing system for multi-source traffic data on highways, comprising:

[0007] The data acquisition standardization module is used to collect multi-source traffic data, preprocess the multi-source traffic data, and obtain a standardized dataset;

[0008] The data source quality assessment module, based on a standardized dataset, evaluates the credibility of each data source and scores each piece of data for timeliness, completeness, and reasonableness, resulting in an enhanced dataset.

[0009] The scenario urgency calculation module, based on the augmented dataset, calculates traffic flow anomaly, vehicle speed anomaly, and weather risk scores, and identifies emergency events to obtain a traffic scenario urgency score.

[0010] The data grading module calculates a comprehensive priority score based on the augmented dataset and the traffic scenario urgency score, and uses a threshold segmentation method to obtain a graded dataset.

[0011] The data classification module extracts multidimensional features from the hierarchical dataset, performs adaptive classification, and obtains a classified dataset.

[0012] The data fusion module, based on the classification dataset, identifies multi-source data with the same traffic parameter and performs conflict detection. It then uses evidence theory to fuse the conflicting data to obtain a fused dataset.

[0013] The results output module, based on the fused dataset, employs a hierarchical classification output and quality feedback mechanism to obtain processing results tailored to different application scenarios.

[0014] In a preferred embodiment, the credibility assessment of each data source includes:

[0015] Based on the historical data accuracy, data reporting stability, and data integrity of the data source, a weighted linear combination method is used to calculate the overall credibility of the data source.

[0016] The credibility of the data source is dynamically updated based on the data verification results. When the data passes verification, a positive incentive update mechanism is used to improve the credibility, and when the data fails verification, a negative penalty update mechanism is used to reduce the credibility.

[0017] In a preferred embodiment, the scoring of timeliness, completeness, and reasonableness for each piece of data includes:

[0018] The timeliness score of the data is calculated based on the difference between the data collection timestamp and the current system time using an exponential decay function.

[0019] Based on the filling status of required and optional fields in the data, a weighted method is used to calculate the data integrity score;

[0020] The reasonableness score of the data is calculated based on the degree of deviation between the data values ​​and the historical statistical distribution.

[0021] In a preferred embodiment, the calculation of the traffic scenario urgency score includes:

[0022] Based on the current traffic flow data of the road segment, the deviation calculation method is used to assess the degree of traffic flow anomaly;

[0023] Based on the vehicle speed data of the current road segment, the segmented deviation calculation method is used to evaluate the degree of vehicle speed abnormality. When the vehicle speed is within the preset normal range, the abnormality is the first value, and when it exceeds the normal range, the abnormality is the relative deviation ratio.

[0024] Meteorological risk scores are calculated based on meteorological monitoring data using a multi-factor weighted method.

[0025] Based on the data text description field, a keyword matching method is used to identify emergency events and build a keyword library containing preset emergency event keywords;

[0026] A hierarchical weighted method is used to calculate the overall urgency of traffic scenarios. When there is an emergency, the urgency is not lower than the first threshold. When there is no emergency but the conventional indicators are severely abnormal, the urgency reaches the second threshold.

[0027] In a preferred embodiment, the comprehensive priority score of the calculated data includes:

[0028] The data integrity score and the reasonableness score are multiplied and combined to obtain the comprehensive data quality score;

[0029] Set weighting coefficients for data source credibility, data timeliness, data quality, and the urgency of the traffic scenario;

[0030] A weighted linear combination method is used to multiply the scores of each dimension by their corresponding weights and then sum them to obtain the comprehensive priority score of the data.

[0031] In a preferred embodiment, the adaptive classification includes:

[0032] Based on the numerical fields of the data, statistical analysis methods are used to extract the statistical characteristics of the data.

[0033] Based on the time series attributes of the data, frequency domain analysis methods are used to extract time series features;

[0034] Spatial features are extracted based on the geographic location information of the data using spatial statistical analysis methods.

[0035] Statistical features, temporal features, and spatial features are combined into a comprehensive feature vector;

[0036] Clustering algorithms are used to classify the comprehensive feature vectors, resulting in a dataset with classification labels.

[0037] In a preferred embodiment, the evidence-based fusion processing of the conflicting data includes:

[0038] Based on timestamps, geographic locations, and data parameter types, multi-source data describing the same traffic parameter are identified and grouped into a data association set.

[0039] Perform a consistency check on multiple data points in a data association set to determine if there are any conflicts.

[0040] For conflicting data, a weighted voting method is used to resolve the data conflict, and a weighted average is calculated based on the credibility of each data source.

[0041] The final fusion result is obtained through weighted fusion calculation.

[0042] In a preferred embodiment, obtaining the hierarchical dataset using the threshold segmentation method includes:

[0043] Multiple priority thresholds are set to divide the overall priority scoring range into different levels;

[0044] Based on the comparison between the comprehensive priority score and the threshold of the data, assign a corresponding priority label to each data point;

[0045] Add a priority level identifier field and a priority score field to each data entry to obtain a dataset with hierarchical identifiers.

[0046] In a preferred embodiment, the calculation of the comprehensive priority score includes:

[0047] Based on the historical accuracy, stability, and completeness of the data source, a weighted linear combination method is used to calculate the data source reputation.

[0048] Based on the time difference between the data collection time and the current time, the data timeliness score is calculated using an exponential decay function.

[0049] Based on the completeness of data fields and the reasonableness of numerical values, a multiplication combination method is used to calculate the data quality score;

[0050] The data source credibility, data timeliness score, data quality score, and traffic scenario urgency are weighted and summed according to set weights to obtain the comprehensive data priority score.

[0051] A computer-readable storage medium for storing computer-readable instructions that, when read by a computer, enable the operation of a highway multi-source traffic data hierarchical classification and processing system.

[0052] The beneficial effects of this invention are as follows:

[0053] By establishing a four-dimensional dynamic grading mechanism based on data source credibility, data timeliness, data integrity, and traffic scenario urgency, the system can comprehensively consider the historical performance of data sources, data quality, and the real-time status of traffic scenarios, dynamically adjusting data processing priorities. Data source credibility is calculated through a weighted combination of historical accuracy, stability, and integrity indicators, and is updated in real time based on the verification results of new data, avoiding interference from low-quality data sources. Data timeliness is quantified using an exponential decay function, ensuring that the latest data receives higher weight. Traffic scenario urgency is determined by integrating traffic flow anomaly, vehicle speed anomaly, weather risk, and emergency event identification results, accurately identifying emergencies requiring immediate response. Compared to traditional static rule-based classification methods, this multi-dimensional dynamic grading mechanism more accurately reflects the actual importance and urgency of data, improving the response speed to emergency events and the targeted nature of data processing.

[0054] By extracting statistical, temporal, and spatial features from the data, a clustering algorithm is used for adaptive classification, and optimal processing strategies are configured for different categories, achieving intelligent and refined data processing. For traffic data with periodic characteristics, the system automatically configures time series analysis processing strategies; for speed data with spatial continuity, the system configures spatial interpolation processing strategies; and for discrete event data, the system configures key information extraction strategies. This adaptive classification mechanism does not require manual pre-definition of data types and processing flows, and can flexibly respond to new data sources and changing data characteristics. Compared with traditional fixed processing flows, it improves the accuracy and adaptability of data processing. Simultaneously, this invention employs a multi-source data fusion method based on evidence theory, assigning a confidence function based on the reputation of each data source, handling data conflicts through Dempster's combination rules, and quantifying the uncertainty of the fusion results. Compared with simple weighted averaging methods, it can more reasonably handle data conflicts and obtain more reliable fusion results. Attached Figure Description

[0055] Figure 1 This is a module diagram of a highway multi-source traffic data hierarchical classification and processing system according to the present invention. Detailed Implementation

[0056] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.

[0057] At least one embodiment of the present invention discloses a hierarchical classification and processing system for multi-source traffic data on highways, such as... Figure 1 As shown, it includes:

[0058] The data acquisition standardization module is used to collect multi-source traffic data, preprocess the multi-source traffic data, and obtain a standardized dataset;

[0059] Specifically, it includes the following:

[0060] Based on various data sources such as high-definition video surveillance cameras, microwave vehicle detectors, ETC gantry systems, meteorological monitoring stations, and vehicle-mounted GPS devices deployed along highways, and using corresponding data acquisition protocols and interfaces, various types of traffic data are acquired.

[0061] Based on high-definition video surveillance cameras deployed along the highway, the system uses the RTSP (Real-Time Streaming Protocol) to access video streams and acquire real-time monitoring video from each camera. For each video stream, metadata information is extracted, including the camera's unique identifier, geographic coordinates (longitude and latitude), video capture timestamp, video resolution, frame rate, and other information, resulting in the raw video dataset. Each record in the raw video dataset contains fields such as video stream address, camera number, location coordinates, and timestamp.

[0062] Based on microwave vehicle detectors deployed along the highway, a data acquisition connection is established using the TCP / IP communication protocol, acquiring detection data from each detector every 30 seconds. For each data packet reported by the detector, information such as traffic flow (number of vehicles passing per unit time), average vehicle speed (unit: kilometers per hour), lane occupancy (unit: percentage), detector number, detection timestamp, and detection point location are parsed to obtain the raw vehicle detector dataset. Each record in the raw vehicle detector dataset contains fields such as detector number, location information, acquisition time, traffic flow value, average vehicle speed value, and lane occupancy value.

[0063] Based on the ETC gantry system deployed along the highway, a dedicated ETC data interface is used to receive vehicle passage records in real time. For each transaction record generated when a vehicle passes through the gantry, information such as the vehicle's ETC card number (anonymized), license plate number (anonymized), vehicle type (classified as passenger car, freight car, etc.), passage timestamp, gantry number, and gantry location are extracted to obtain the original ETC dataset. Each record in the original ETC dataset contains fields such as anonymized vehicle identifier, vehicle type code, passage time, gantry number, and location.

[0064] Based on meteorological monitoring stations deployed along the highway, meteorological observation data is obtained every 10 minutes using an HTTP API interface. For each meteorological station's reported data, information such as temperature (degrees Celsius), relative humidity (percentage), visibility (meters), precipitation (millimeters), wind speed (meters per second), wind direction, meteorological station number, observation time, and meteorological station location are extracted to obtain the raw meteorological dataset. Each record in the raw meteorological dataset contains the meteorological station number, location, observation time, and values ​​for various meteorological parameters.

[0065] Based on in-vehicle GPS devices connected in cooperation with commercial vehicles, the location information of approximately 100,000 vehicles is received in real time via a GPS data reporting interface. For each vehicle's GPS data packet reported every 10 seconds, information such as vehicle identifier (anonymized), current location's longitude and latitude coordinates, instantaneous speed (in kilometers per hour), driving direction angle, GPS positioning timestamp, and positioning accuracy are extracted to obtain the raw GPS trajectory dataset. Each record in the raw GPS trajectory dataset contains fields such as anonymized vehicle identifier, longitude and latitude coordinates, speed, direction, timestamp, and positioning accuracy.

[0066] Based on the raw video dataset, raw vehicle detector dataset, raw ETC dataset, raw weather dataset, and raw GPS trajectory dataset obtained above, a unified JSON format conversion method was used for format standardization. For each type of raw data, it was converted into a standard JSON object containing the following unified fields: a data type identifier field (values ​​can be one of video, detector, ETC, weather, or GPS), a unique data source number field, a data collection timestamp field (uniformly converted to UTC time format), a geographic location field (including longitude and latitude subfields), a data content field (storing specific numerical or descriptive information for this type of data), and a metadata field (storing additional information such as data resolution and collection frequency), resulting in a standardized dataset. All data in the standardized dataset uses the same JSON format and a unified time base, facilitating subsequent unified processing.

[0067] The output of this module is a standardized dataset, which will serve as the input for the next module.

[0068] The data source quality assessment module, based on a standardized dataset, evaluates the credibility of each data source and scores each piece of data for timeliness, completeness, and reasonableness, resulting in an enhanced dataset.

[0069] Specifically, it includes the following:

[0070] Based on the historical reporting records of each data source stored in the historical data quality record library, for each data source, the sliding time window method is used to select historical data of the past 7 days as the evaluation sample. Each data report of the data source within the time window is compared with the real value confirmed by multi-source verification. The number of data reports with an acceptable error between the reported data and the real value is recorded as the number of correct data reports. The total number of data reports of the data source within the time window is recorded as the total number of data reports.

[0071] The accuracy metric for this data source is calculated using the following steps: First, obtain the number of correctly reported data entries. Extract the number of correctly reported data entries from the statistical results within the time window. Second, obtain the total number of data entries. Extract the total number of data entries reported by the data source within the time window from the statistical results. Third, perform a division operation, dividing the number of correctly reported data entries by the total number of data entries. The result of this division operation is the accuracy metric for the data source. The accuracy metric ranges from 0 to 1; a value closer to 1 indicates higher accuracy.

[0072] Furthermore, a weighted accuracy calculation method can be used instead of a simple accuracy calculation to more comprehensively reflect the historical performance of the data source. Specifically, the historical data within the time window is divided into multiple sub-time periods according to time sequence. Sub-time periods closer to the current moment are assigned higher weight coefficients, reflecting the greater impact of the recent performance of the data source on the current credibility. An exponential decay weight function is used to calculate the weight of each sub-time period. For the k-th sub-time period (k ranges from 1 to M, with larger k indicating closer to the current moment), its weight is the value of the exponential decay function, where the decay coefficient is usually taken as 0.2. The weighted accuracy index is calculated using the following steps: Calculate the sum of weighted accuracy products. Iterate through all sub-time periods. For each sub-time period k, multiply its weight coefficient by the accuracy of that sub-time period. Sum the products of all sub-time periods to obtain the numerator value. Calculate the sum of weight coefficients. Sum the weight coefficients of all sub-time periods to obtain the denominator value. Perform a division operation, dividing the numerator value by the denominator value. The weighted accuracy index is obtained, and the result of the division operation is the weighted accuracy index of the data source. Where M represents the total number of sub-time periods, and k represents the sequence number of the sub-time period. This method makes reputation assessment pay more attention to the recent performance of the data source and respond promptly to changes in the quality of the data source.

[0073] Based on data reporting records from the same data source within a time window, statistical analysis methods are used to evaluate the stability of data reporting for each data source. The specific steps are as follows: Extract the time interval between all two adjacent data reports from the data source within the time window; calculate the standard deviation of these time intervals, denoted as the time interval standard deviation; the smaller the standard deviation, the more regular and stable the reporting time. For numerical data (such as vehicle speed, traffic flow, etc.), the stability of the data values ​​needs to be evaluated: calculate the coefficient of variation of the reported data values ​​from the data source. The coefficient of variation equals the standard deviation of the data values ​​divided by the mean, denoted as the data value coefficient of variation; the smaller the coefficient of variation, the more stable the data values. Based on the reporting time interval standard deviation and the data value coefficient of variation, the following steps are used to calculate the stability index of the data source: add the time interval standard deviation of the data source to the data value coefficient of variation to obtain the sum of instability values; add the sum of instability values ​​to a constant 1 to obtain the denominator value; calculate the reciprocal of the denominator value, i.e., divide the denominator value by the constant 1; the result of the reciprocal operation is the stability index of the data source. The stability index ranges from 0 to 1 (excluding 0), with values ​​closer to 1 indicating better data source stability.

[0074] Based on the data records reported by the data sources, for each data source, a data missing rate statistical method is used to assess its data integrity. According to the normal reporting frequency of this type of data source, the expected number of data reports within a time window is calculated and recorded as the expected reporting number (e.g., a detector that reports once every 30 seconds is expected to report 20,160 times in 7 days). The actual number of data reports from this data source is counted and recorded as the actual reporting number. The number of missing data is calculated as the expected reporting number minus the actual reporting number. For the actually reported data, the required fields (such as timestamp, location, and main numeric fields) of each data entry are checked for missing data. For complete data entry, the number of data entries with missing required fields is recorded as the field missing count. Based on the number of missing data entries and the number of missing fields, the following steps are used to calculate the data source integrity index: Calculate the field missing penalty value by multiplying the number of missing data entries in the required fields by a weighting coefficient of 2; calculate the total missing penalty value by adding the number of missing data entries to the field missing penalty value; calculate the missing rate by dividing the total missing penalty value by the expected number of data reports; and calculate the integrity index by subtracting the missing rate from a constant of 1 to obtain the data source integrity index.

[0075] The integrity metric for this data source is calculated based on the number of missing data entries and the number of missing fields, using the following steps: Multiply the number of data entries with missing required fields by a weighting factor of 2 to obtain the field missing penalty value; add the number of missing data entries to the field missing penalty value to obtain the total missing penalty value; divide the total missing penalty value by the expected number of data reports to obtain the missing rate; subtract the missing rate from the constant 1 to obtain the integrity metric for this data source. The integrity metric assigns double weight to field missing data because field missing data has a greater impact on subsequent processing than missing entire data entries. Theoretically, the integrity metric ranges from negative infinity to 1, but in practice, negative values ​​are truncated to 0.

[0076] Based on the accuracy, stability, and integrity indices obtained from the above calculations, a weighted linear combination method is used to calculate the overall credibility of the data source.

[0077] Data preprocessing steps: Check the value range of the three indicators to ensure they are all normalized to the 0-1 range. If any indicator's value exceeds this range, use the minimum-maximum normalization method to rescale it to the 0-1 range. For potential outliers (such as negative values ​​due to calculation errors or values ​​exceeding 1), truncation is performed, setting values ​​less than 0 to 0 and values ​​greater than 1 to 1. Based on the importance of the three dimensions—accuracy, stability, and completeness—set weight coefficients: accuracy weight 0.5 (most important), stability weight 0.3 (second most important), and completeness weight 0.2. The three weights must satisfy the normalization condition (the sum of the weights equals 1). Reputation calculation steps: Calculate the accuracy contribution value by multiplying the accuracy indicator by its weight; calculate the stability contribution value by multiplying the stability indicator by its weight; calculate the completeness contribution value by multiplying the completeness indicator by its weight; calculate the overall reputation score by adding the accuracy contribution value, the stability contribution value, and the completeness contribution value. The overall reputation score ranges from 0 to 1, with values ​​closer to 1 indicating a more reliable data source.

[0078] Based on the real-time performance of the data source, an incremental learning method is used to dynamically update the reputation score. For each newly received and subsequently validated piece of data, the reputation score of its respective data source is adjusted according to the validation results. When data is confirmed as accurate through multi-source cross-validation, a positive incentive update mechanism is used, updating the data source's reputation score according to the following steps: subtract the current reputation score of the data source from a constant of 1 to obtain the reputation score improvement space; multiply the reputation score improvement space by the positive incentive coefficient (usually 0.01) to obtain the reputation score increment; add the current reputation score of the data source to the reputation score increment to obtain the updated reputation score. When data is confirmed as erroneous by cross-validation, a negative penalty update mechanism is used, updating the data source's reputation score according to the following steps: multiply the current reputation score of the data source by the negative penalty coefficient (usually 0.05) to obtain the reputation score penalty amount; subtract the reputation score penalty amount from the current reputation score of the data source to obtain the updated reputation score. The positive incentive coefficient is usually 0.01, and the negative penalty coefficient is usually 0.05. The negative penalty coefficient being greater than the positive incentive coefficient reflects a stricter constraint on erroneous data. Through continuous dynamic updates, the data source reputation can reflect its latest working status, resulting in a real-time reputation score set for each data source (including the reputation of data source 1, the reputation of data source 2, and the reputation of data source n), where n is the total number of data sources.

[0079] Based on the collection timestamp and current system time of each data point in the standardized dataset, an exponential decay function is used to quantify the timeliness of the data. For each data point in the standardized dataset, its collection timestamp is extracted, the current system timestamp is obtained, and the delay time from data collection to the present is calculated. The time decay coefficient is determined according to the data type. For emergency event data (such as accident alarms and anomaly detection results), a larger decay coefficient of 0.1 seconds to the power of negative 1 is set; for regular traffic statistics data, a smaller decay coefficient of 0.01 seconds to the power of negative 1 is set; and for slowly changing data such as meteorological data, an even smaller decay coefficient of 0.001 seconds to the power of negative 1 is set.

[0080] The timeliness score of the data is calculated using the following steps: Subtract the data collection timestamp from the current system timestamp to obtain the data delay time; multiply the time decay coefficient by the data delay time to obtain the decay exponent; calculate the timeliness score by exponentially multiplying the natural constant e to the base and the negative decay exponent to obtain the timeliness score for that data entry. The timeliness score ranges from 0 to 1. Newly collected data has a timeliness score of 1, and the score decays exponentially over time.

[0081] Based on the field filling status of each data point in the standardized dataset, a weighted field completeness method is used to assess data integrity. For each data point, a list of required fields and a list of optional fields are determined according to its data type. For example, for vehicle detector data, required fields include detector number, timestamp, location coordinates, traffic flow, and vehicle speed, while optional fields include lane occupancy rate and large vehicle percentage. All fields of the data point are traversed, and the total number of required fields and the number of required fields actually filled (the values ​​are not empty and are not missing) are counted. The number of missing required fields is calculated (equal to the total number of required fields minus the number actually filled). Similarly, the total number of optional fields and the number of missing optional fields are counted.

[0082] The weighted completeness score is calculated using the following steps: Divide the number of required fields actually filled (i.e., the total number of required fields minus the number of missing required fields) by the total number of required fields to obtain the required field completeness rate; divide the number of optional fields actually filled (i.e., the total number of optional fields minus the number of missing optional fields) by the total number of optional fields to obtain the optional field completeness rate; multiply the required field completeness rate by a weight of 0.8 to obtain the contribution value of required fields to the completeness score; multiply the optional field completeness rate by a weight of 0.2 to obtain the contribution value of optional fields to the completeness score; add the contribution values ​​of required fields and optional fields to obtain the completeness score of the data.

[0083] Based on the numerical fields of each data point in the standardized dataset, the historical statistical range test method is used to evaluate the reasonableness of the data. For the main numerical fields of each data point (such as vehicle speed, traffic flow, etc.), the historical statistical range of this type of data under similar scenarios (same road segment, same time period) is queried from the historical data statistics database. The historical minimum value, historical maximum value and reasonable fluctuation range (usually set to 10% of the historical range, that is, the reasonable fluctuation range is equal to 0.1 multiplied by the difference between the historical maximum value and the historical minimum value) are obtained. The value of the current data is extracted and compared with the historical statistical range.

[0084] The following segmented scoring rules are used to calculate the reasonableness score: Determine the numerical range and examine the relationship between the current data value and the historical statistical range; apply the segmented scoring rules: if the value is within the historical range (historical minimum ≤ current value ≤ historical maximum value), the reasonableness score is 1; if the value is within the tolerance deviation range (historical minimum minus reasonable fluctuation range < current value < historical minimum, or historical maximum < current value < historical maximum plus reasonable fluctuation range), the reasonableness score is 0.5; if the value exceeds the tolerance range (current value ≤ historical minimum minus reasonable fluctuation range, or current value ≥ historical maximum plus reasonable fluctuation range), the reasonableness score is 0. Here, historical minimum and historical maximum represent the boundaries of the historical statistical range, and reasonable fluctuation range represents the tolerable deviation range. A score of 1 is given when the value is within the historical range, 0.5 is given when it is within the tolerable deviation range, and 0 is given when it exceeds the tolerable range, indicating a serious data anomaly.

[0085] The output of this module is an augmented dataset with quality annotations. Each data point includes quality annotation information such as data source reputation, timeliness score, completeness score, and reasonableness score, which serves as the input for the next module.

[0086] Furthermore, a machine learning-based reputation prediction model can be used to replace the simple weighted linear combination method to capture the non-linear relationship between accuracy, stability, and integrity. Specifically, historical performance data from a large number of data sources are collected as training samples. Each sample contains accuracy, stability, and integrity indicators at a certain moment as features, and the actual reliability performance of the data source in the future as a label. A gradient boosting decision tree model is used for training, and the model can automatically learn the interaction and non-linear relationship between the indicators. In practical applications, the calculated accuracy, stability, and integrity metrics are input into the trained model, and the predicted reputation score is obtained by following these steps: Data preprocessing is performed, standardizing the accuracy (accuracy score of the data source), stability (stability score of the data source), and integrity (integrity score of the data source) to a mean of 0 and a standard deviation of 1, eliminating the influence of differences in the scale of different metrics; Model prediction is performed, inputting the three standardized metrics into the trained gradient boosting decision tree model (GBDT model function), which contains 100 weak learners (decision trees), each with a maximum depth of 6 and a learning rate of 0.1; Output mapping is performed, mapping the continuous values ​​of the model output to the interval between 0 and 1 using the sigmoid function, specifically by dividing 1 by (1 plus the natural constant e raised to the power of negative x), where x is the original output value of the model; The final reputation score is obtained, and the result after mapping using the sigmoid function is the machine learning predicted reputation score of data source i (the machine learning reputation score of the data source).

[0087] Furthermore, a context-aware dynamic thresholding method can be used to replace the fixed historical statistical range to adapt to the dynamic changes in traffic data. Specifically, it considers not only historical statistical data from the same period but also current contextual factors such as traffic events, weather conditions, and holidays. A multi-dimensional historical data index is constructed, and the distribution of historical data under the most similar scenarios is retrieved based on the current contextual conditions (such as rainy days or weekday morning rush hour), dynamically determining the threshold range for the reasonableness score.

[0088] A contextual similarity calculation method is adopted, following these steps: Weather similarity is multiplied by a weighting coefficient to obtain the weather similarity contribution value; time similarity is multiplied by a weighting coefficient to obtain the time similarity contribution value; event similarity is multiplied by a weighting coefficient to obtain the event similarity contribution value; the three similarity contribution values ​​(weather, time, and event) are summed to obtain the comprehensive contextual similarity. The reasonableness scoring threshold is dynamically adjusted based on the historical scenario with the highest similarity. This method improves the adaptability of data reasonableness assessment, reduces the misjudgment rate, and can promptly adjust the assessment criteria, especially when significant changes occur in traffic conditions.

[0089] The scenario urgency calculation module, based on the augmented dataset, calculates traffic flow anomaly, vehicle speed anomaly, and weather risk scores, and identifies emergency events to obtain a traffic scenario urgency score.

[0090] Specifically, it includes the following:

[0091] Based on real-time traffic flow data and historical traffic flow statistical models for the current road segment, a deviation calculation method is used to assess the degree of traffic flow anomaly. For each road segment, the current traffic flow observation value is extracted from the enhanced dataset with quality annotations. The historical average traffic flow and historical traffic flow standard deviation for the same time period (such as weekday morning rush hour, weekend afternoon, etc.) are queried from the historical traffic flow database. The traffic flow deviation value is calculated by subtracting the historical average traffic flow for the same period from the current traffic flow observation value and then taking the absolute value. A division-by-zero protection constant is set. To avoid division-by-zero errors, a small constant equal to 1 is set as the denominator protection value. The denominator value is calculated by adding the historical average traffic flow for the same period to the small constant. A division operation is performed to divide the traffic flow deviation value by the denominator value to obtain the traffic flow anomaly degree for the road segment. The traffic flow anomaly degree ranges from 0 to positive infinity, with a larger value indicating a more severe deviation from the normal traffic flow state.

[0092] Based on the real-time vehicle speed data and normal speed range model of the current road segment, a segmented deviation calculation method is used to assess the degree of vehicle speed anomalies. For each road segment, the average vehicle speed observation value (unit: kilometers per hour) at the current moment is extracted from the augmented dataset with quality annotation, and the normal speed range of the road segment is obtained from the road segment feature database (usually determined based on the design speed and historical speed distribution of the road segment). The following segmented calculation rules are used to obtain the vehicle speed anomaly degree: Determine the vehicle speed range and check the relationship between the current average vehicle speed and the normal vehicle speed range; calculate the anomaly degree segment by segment: If the vehicle speed is within the normal range (lower bound of the normal speed range ≤ current average vehicle speed ≤ upper bound of the normal speed range), the vehicle speed anomaly degree is equal to 0; if the vehicle speed is lower than the lower bound of the normal range (current average vehicle speed is less than the lower bound of the normal speed range), calculate the relative deviation ratio by subtracting the current vehicle speed from the lower bound of the normal range, and then dividing by the lower bound of the normal range to obtain the vehicle speed anomaly degree; if the vehicle speed is higher than the upper bound of the normal range (current average vehicle speed is greater than the upper bound of the normal speed range), calculate the relative deviation ratio by subtracting the upper bound of the normal range from the current vehicle speed, and then dividing by the upper bound of the normal range to obtain the vehicle speed anomaly degree. Specifically, the anomaly degree is 0 when the vehicle speed is within the normal range, and the anomaly degree is the relative deviation ratio when the speed is lower or higher than the normal range. The value range of the vehicle speed anomaly degree is from 0 to positive infinity.

[0093] Based on current meteorological observation data, a multi-factor weighted method is used to calculate the environmental risk index. Current observation data from the nearest meteorological station to the road segment is extracted from a quality-annotated enhanced dataset, including visibility, precipitation intensity, and road icing probability (calculated based on temperature and humidity, with values ​​ranging from 0 to 1). A safe threshold of 10,000 meters is set for visibility (visibility exceeding this value is considered to have no impact on traffic), and a danger threshold of 50 millimeters per hour (heavy rain level) is set for precipitation intensity. Weighting coefficients are assigned according to the degree of influence of meteorological factors on traffic safety: visibility weight is 0.4, precipitation weight is 0.4, icing weight is 0.2, and the sum of the weights is 1. The meteorological risk score for this road section is calculated by: calculating the visibility risk contribution (dividing the current visibility by the safety threshold, subtracting the ratio from 1 using a constant, and then multiplying by the visibility weight); calculating the precipitation risk contribution (dividing the current precipitation intensity by the danger threshold and then multiplying by the precipitation weight); calculating the icing risk contribution (multiplying the road surface icing probability by the icing weight); and finally calculating the total meteorological risk (summing the visibility risk contribution, precipitation risk contribution, and icing risk contribution). The meteorological risk score ranges from 0 to 1, with higher values ​​indicating more severe weather conditions.

[0094] Based on the text information of the data content in the augmented dataset with quality annotation, a keyword matching method is used to identify emergency events. For the text description field of each data point (such as event detection results in video surveillance, text descriptions reported manually, etc.), an emergency event keyword library is constructed, including keywords such as "accident", "collision", "fire", "explosion", "personnel casualties", "vehicle rollover", and "hazardous material leakage". The data content is segmented and matched with keywords. If any emergency event keyword is matched, the emergency event identifier associated with the data is set to 1; otherwise, the emergency event identifier is set to 0. The emergency event identifier is a binary variable that directly indicates whether there is an emergency that requires immediate response.

[0095] Based on the calculated traffic flow anomaly, vehicle speed anomaly, meteorological risk score, and emergency event identifier, a hierarchical weighted method is used to calculate the comprehensive traffic scenario urgency of the road segment. This includes the following steps: Since the traffic flow anomaly ranges from 0 to positive infinity, a logistic function transformation is used to map it to the 0-1 interval. The normalized traffic flow anomaly is calculated by dividing it by (traffic flow anomaly plus 1). The normalized vehicle speed anomaly is also calculated by using a logistic function transformation, dividing it by (vehicle speed anomaly plus 1). The meteorological risk score is maintained, as its range is already 0-1 and no transformation is needed. The average value of regular anomalies is calculated by adding the normalized traffic flow anomaly, the normalized vehicle speed anomaly, and the meteorological risk score, and then dividing by 3. Finally, the comprehensive urgency is calculated by multiplying the emergency event identifier by a weight of 0.5 and the average value of regular anomalies by a weight of 0.5, and then adding the two results to obtain the comprehensive traffic scenario urgency of the road segment. Specifically, when an emergency occurs, the urgency level is at least 0.5 even if the general indicators are normal; when there is no emergency but the general indicators are severely abnormal, the urgency level can still reach 0.5. This yields a set of traffic scenario urgency scores for all road segments, containing the urgency scores for all road segments.

[0096] The output of this module is a set of traffic scenario urgency scores for each road segment, which will serve as the input for the next module.

[0097] Furthermore, a time-series anomaly detection method based on recurrent neural networks can be used to replace the anomaly calculation based on statistical thresholds, in order to more accurately capture the time-series evolution patterns of traffic conditions. Specifically, a Long Short-Term Memory (LSTM) network is used to model and train the historical traffic flow and vehicle speed time series for each road segment, and the model learns the time-series variation patterns of normal traffic flow.

[0098] The core of an LSTM network lies in its gating mechanism, including the forget gate, input gate, and output gate. The forget gate determines what information is discarded from the cell state. It concatenates the previous hidden state and the current input into a vector, multiplies it by the forget gate weight matrix, adds a bias, and finally passes it through the sigmoid activation function to obtain the forget gate output. The input gate determines what new information is stored in the cell state. It concatenates the previous hidden state and the current input into a vector, multiplies it by the input gate weight matrix, adds a bias, and finally passes it through the sigmoid activation function to obtain the input gate output. Alternatively, it concatenates the previous hidden state and the current input into a vector, multiplies it by the candidate value weight matrix, adds a bias, and so on. Finally, the candidate value vector is obtained through the tanh activation function. The forget gate output is multiplied element-wise by the cell state at the previous time step, and then the input gate output is multiplied element-wise by the candidate value vector. The two results are added together to obtain the cell state at the current time step. The output gate determines the output value. The hidden state at the previous time step and the current input are concatenated into a vector, multiplied by the output gate weight matrix, a bias is added, and finally, the output gate output is obtained through the sigmoid activation function. The output gate output is then multiplied element-wise by the result of the current cell state after the tanh activation function to obtain the hidden state at the current time step. Here, the forget gate, input gate, and output gate are the activation values ​​of each gating mechanism, and the sigmoid function is the activation function.

[0099] In real-time applications, the most recent observation data of the road segment is input into the trained model. The model outputs a predicted value for the traffic state at the next moment, and the deviation between the actual observation value and the model prediction value is calculated as the anomaly index. The process involves acquiring actual traffic state observation values ​​of the road segment from sensors; inputting historical time series data into the trained LSTM model to obtain predicted values; calculating the prediction standard deviation based on historical prediction errors; and dividing the absolute difference between the actual observation value and the predicted value by the prediction standard deviation to obtain the LSTM-based anomaly index for the road segment.

[0100] The data grading module calculates a comprehensive priority score based on the augmented dataset and the traffic scenario urgency score, and uses a threshold segmentation method to obtain a graded dataset.

[0101] Specifically, it includes the following:

[0102] Based on each data point in the quality-annotated augmented dataset, the association between the data and various scoring indicators is established. For each data point in the quality-annotated augmented dataset, its data source is determined by its unique data source ID field, and the credibility score of that data source is obtained. The timeliness score, completeness score, and reasonableness score of the data point itself are obtained. The associated road segment is determined by the data's geographic location field (by matching geographic location with road segment range), and the traffic scenario urgency score of that road segment is obtained. These five scoring indicators are associated with the data to obtain the complete score set of the data.

[0103] Based on the data association score set, a multi-dimensional weighted linear combination method is used to calculate the comprehensive priority score of the data. Before weighting and combining, data preprocessing is required for each score indicator to ensure the comparability of indicators with different dimensions: check the value range of all score indicators to ensure that they have been normalized to the 0 to 1 interval; for possible outliers, truncation is used to limit values ​​outside the range to the 0 to 1 interval; since different score indicators may have different distribution characteristics (such as some indicator values ​​being generally high or low), standardization can be selectively used (subtract the mean from the original value and divide by the standard deviation) before remapping to the 0 to 1 interval to balance the influence of each indicator.

[0104] Since both completeness and reasonableness scores reflect the quality of the data itself, they are combined by multiplication to form a comprehensive data quality score (completeness score multiplied by reasonableness score). This multiplicative combination ensures that a high-quality score is obtained only when the data is both complete and reasonable. Four weight coefficients are set for each dimension: credibility weight, timeliness weight, data quality weight, and scenario urgency weight. Based on application requirements, scenario urgency has the greatest impact on priority, so its weight is set to 0.4. The remaining weights are evenly distributed among the other three dimensions, with credibility weight set to 0.2, timeliness weight set to 0.2, and data quality weight set to 0.2, satisfying the normalization condition (the sum of all weights equals 1).

[0105] A weighted linear combination method is used to calculate the comprehensive priority score of the data. Based on the reputation contribution, timeliness contribution, data quality contribution, and urgency contribution calculated above, the comprehensive priority score is calculated as follows: Reputation contribution is calculated by multiplying the data source reputation score by a reputation weight; timeliness contribution is calculated by multiplying the data timeliness score by a timeliness weight; data quality contribution is calculated by multiplying the completeness score and the reasonableness score, and then multiplying by the data quality weight; urgency contribution is calculated by multiplying the traffic scenario urgency by a scenario urgency weight; and the comprehensive priority score is calculated by summing the four contribution values. The comprehensive priority score ranges from 0 to 1, with a higher value indicating more important and urgent data that should be processed first.

[0106] Based on the calculated comprehensive priority score, a threshold segmentation method was used to divide the data into four priority levels. According to the processing requirements and resource allocation strategies for different priority data, three tiered thresholds were set: Level 1 threshold was 0.8, Level 2 threshold was 0.6, and Level 3 threshold was 0.4.

[0107] For each piece of data, a segmented judgment rule is used to determine its priority level based on the comprehensive priority score: when the comprehensive priority score is greater than or equal to 0.8, it is judged as Level 1 (urgent) priority, corresponding to data such as traffic accidents and emergency events that need to be handled immediately; when the comprehensive priority score is greater than or equal to 0.6 and less than 0.8, it is judged as Level 2 (important) priority, corresponding to data such as traffic anomalies and congestion warnings that need to be handled with priority; when the comprehensive priority score is greater than or equal to 0.4 and less than 0.6, it is judged as Level 3 (normal) priority, corresponding to data such as normal traffic flow and routine weather data; when the comprehensive priority score is less than 0.4, it is judged as Level 4 (low priority), corresponding to data from low-quality data sources and data that has seriously timed out.

[0108] Add a priority level identifier field (with values ​​of 1, 2, 3, and 4) and a priority score field (to store the overall priority score value) to each data entry to obtain a dataset with graded identifiers.

[0109] The output of this module is a dataset with hierarchical labels, where each data entry contains a priority level label field and a priority score field, which will serve as the input for the next module.

[0110] Furthermore, a dynamic threshold adjustment mechanism can be used to replace the fixed hierarchical thresholds to adapt to dynamic changes in system load. Specifically, the length and processing latency of each priority data queue in the current system are monitored. When the first-priority data queue is severely backlogged, the first-priority threshold is appropriately increased, so that only more urgent data is classified as first-priority, thus reducing the pressure on the first-priority queue.

[0111] Dynamic threshold adjustment is achieved through primary and secondary threshold adjustments. The primary threshold adjustment calculation steps are as follows: calculate the primary queue load ratio by dividing the current primary queue length by its capacity; multiply the load ratio by an adjustment coefficient to obtain the threshold adjustment amount; and add the base primary threshold to the threshold adjustment amount to obtain the new primary threshold. The secondary threshold adjustment is as follows: calculate the secondary data latency ratio by dividing the average secondary data processing latency by the secondary data latency threshold to obtain the latency ratio; multiply the latency ratio by an adjustment coefficient to obtain the threshold adjustment amount; and add the base secondary threshold to the threshold adjustment amount to obtain the new secondary threshold. When the overall system load is low, the thresholds at each level are appropriately reduced, allowing more data to receive higher priority processing and improving the timeliness of data processing. This method improves the system's adaptability and resource utilization efficiency, maximizing data processing throughput while ensuring timely processing of urgent data.

[0112] The data classification module extracts multidimensional features from the hierarchical dataset, performs adaptive classification, and obtains a classified dataset.

[0113] Specifically, it includes the following:

[0114] Based on the numerical fields in the hierarchically labeled dataset, statistical analysis methods are used to extract the statistical features of the data. For each data point, all numerical subfields (such as traffic flow, vehicle speed, temperature, etc.) are extracted from its data content field. For individual values, a global normalization method is used to scale them to a standard range to facilitate comparison of data with different dimensions. For data with short-term historical sequences (such as the most recent 10 observations), the mean, standard deviation, skewness coefficient, and kurtosis coefficient of the sequence are calculated. The mean represents the central tendency of the data, the standard deviation represents the dispersion of the data, the skewness coefficient represents the asymmetry of the distribution (positive skewness indicates a long right tail, negative skewness indicates a long left tail), and the kurtosis coefficient represents the sharpness of the distribution (larger kurtosis indicates a more concentrated distribution). These statistics are organized into a statistical feature vector, which comprehensively describes the numerical distribution characteristics of the data. Statistical features can reflect the numerical distribution characteristics of data, and data with similar statistical features often have similar processing requirements.

[0115] Based on the time dimension information of the data, frequency domain analysis is used to extract time-series features. For data types with time-series attributes (such as continuous traffic flow data from vehicle inspection devices and continuous temperature data from weather stations), historical observation sequences from the same data source within the most recent time window are collected. A Fast Fourier Transform (FFT) is performed on this time series to convert the time-domain signal to the frequency domain, identifying the main periodic components in the signal. Extracted time-series features include: main period (the period length of the frequency component with the highest energy in the spectrum), period intensity (the amplitude of this frequency component, reflecting the significance of periodicity), trend slope (the slope parameter obtained by linear regression fitting of the time series; positive values ​​indicate an upward trend, negative values ​​indicate a downward trend, and the absolute value indicates the trend intensity), and stationarity index (determining whether the sequence is stationary through unit root tests; the statistical characteristics of a stationary sequence do not change over time). These features are organized into a time-series feature vector, which reflects the pattern of data evolution over time. Time-series features can reflect the changing patterns of data over time and are instructive for selecting appropriate time-series analysis methods.

[0116] Furthermore, wavelet transform can be used instead of fast Fourier transform for time series feature extraction to better capture the time-frequency characteristics of non-stationary signals. Specifically, a suitable mother wavelet function (such as Daubechies wavelet or Morlet wavelet) is selected to perform multi-scale wavelet decomposition on the time series to obtain wavelet coefficients at different scales.

[0117] Wavelet transform is used to analyze the characteristics of signals at different time scales. The steps include: selecting a suitable mother wavelet function, such as the Morlet wavelet or Daubechies wavelet; setting scale and translation parameters, where the scale parameter controls the scaling of the wavelet and the translation parameter controls its translation; performing an inner product operation between the original signal and the complex conjugate of the scale-scaled and time-translated mother wavelet function, and normalizing by dividing by the square root of the scale parameter to obtain the wavelet coefficients; iterating through all scales and translations, repeating the above calculations for all scales and translation parameters of interest to obtain the complete wavelet transform result.

[0118] This method extracts features such as energy distribution, wavelet entropy, and singularity index from wavelet coefficients. For the j-th scale, the total energy of that scale is obtained by summing the squares of the moduli of all wavelet coefficients at that scale. The energy at the j-th scale is then divided by the sum of the energies of all scales to obtain the energy probability distribution. The energy probability at each scale is multiplied by its logarithm, and the sum is taken over all scales and then negative to obtain the wavelet entropy. Compared with Fourier transform, wavelet transform has better time-frequency localization characteristics, making it particularly suitable for analyzing common sudden events and non-stationary processes in traffic data. This method improves the ability of time-series feature extraction to express complex time-varying signals, resulting in more accurate data classification.

[0119] Based on the geographic location information of the data, spatial statistical analysis methods are used to extract spatial features. For data with geographic coordinates, the spatial location of data points is determined according to their geographic location field. Other data points that are close in time (within the same time window) and spatially adjacent (within a certain distance range) are collected to construct a spatial neighborhood dataset. The spatial autocorrelation coefficient between the value of the current data point and the values ​​of other data points in the spatial neighborhood is calculated, and Moran's I index is used to measure spatial clustering. The meaning of Moran's I index is: when Moran's I index is close to 1, it indicates positive spatial autocorrelation (adjacent positions have similar values, showing spatial clustering); when Moran's I index is close to -1, it indicates negative spatial autocorrelation (adjacent positions have opposite values, showing a checkerboard distribution); when Moran's I index is close to 0, it indicates spatial random distribution (no obvious spatial pattern). The degree of spatial variability, i.e., the variance of data values ​​in the neighborhood, is calculated to reflect the spatial dispersion of the data. The Moran's I index and the degree of spatial variability are organized into a spatial feature vector, which reflects the spatial distribution characteristics of the data.

[0120] Based on the extracted statistical, temporal, and spatial feature vectors, a comprehensive feature vector is constructed using feature concatenation and dimensionality reduction methods. The three types of feature vectors are concatenated sequentially to form a high-dimensional original feature vector. Since the dimensions and numerical ranges of different feature types may vary significantly, standardization is employed to ensure the comparability of features across dimensions. For the high-dimensional feature vector, Principal Component Analysis (PCA) is used for dimensionality reduction, retaining principal components with a cumulative variance contribution rate of 95%, thus preserving the main information of the features while reducing computational complexity. This yields a comprehensive feature vector for each data point, which fully characterizes the data's properties in the statistical, temporal, and spatial dimensions.

[0121] Based on the comprehensive feature vector, the K-means clustering algorithm is used for adaptive data classification. The number of clusters K is determined, and the Elbow Method is used to select the optimal number of clusters. By calculating the sum of squares within each cluster (WCSS) under different K values, the K value where the WCSS decreasing trend shows a clear inflection point is selected as the optimal number of clusters. The K-means clustering algorithm is executed, dividing all data into K categories according to their comprehensive feature vector. Each clustering result is analyzed, and the priority distribution, data type distribution, spatiotemporal distribution, and other characteristics of the data in each cluster are statistically analyzed to determine the typical characteristics of each cluster and the recommended processing strategy.

[0122] Based on the clustering results and feature analysis of each cluster, a corresponding processing strategy is matched for each data category. For clusters containing a large amount of urgent data, a real-time processing strategy is matched, using a streaming computing framework for millisecond-level response; for clusters containing a large amount of time-series data, a time-series analysis strategy is matched, using an ARIMA model or LSTM network for predictive analysis; for clusters containing a large amount of spatially related data, a spatial analysis strategy is matched, using a Geographic Information System (GIS) for spatial interpolation and hotspot analysis; for clusters containing a large amount of low-quality data, a data cleaning strategy is matched, using anomaly detection and data repair algorithms for quality improvement. A category identifier field `Category_id` (representing the cluster number) and a processing strategy field `Processing_strategy` (representing the recommended processing method) are added to each data entry, resulting in a dataset with category identifiers and processing strategies.

[0123] The output of this module is a dataset with category identifiers and processing strategies. Each data entry contains the Category_id and Processing_strategy fields, which will serve as the input for the next module.

[0124] The data fusion module, based on the classification dataset, identifies multi-source data with the same traffic parameter and performs conflict detection. The conflict data is then fused using evidence theory to obtain a fused dataset.

[0125] Specifically, it includes the following:

[0126] Based on multiple data sources with the same data type and similar geographical locations, a consistency check method is used to verify the reliability of the data. For each data point, based on its data type (data_type field) and geographical location (location field), observation data from similar data sources that are close in time (time difference less than 5 minutes) and spatially adjacent (distance less than 1 kilometer) are searched to construct a validation dataset. For numerical data (such as vehicle speed and traffic flow), the numerical difference between the current data and each data point in the validation dataset is calculated using the following consistency score calculation method: Calculate the mean absolute deviation (MAD) by averaging the absolute values ​​of the differences between the current data value and the values ​​of each data point in the validation dataset; set a consistency threshold by defining an acceptable deviation range based on the data type (e.g., a threshold of 10 kilometers per hour for vehicle speed data); calculate the consistency score. If the MAD is less than the consistency threshold, the consistency score is 1; otherwise, the consistency score is the threshold divided by (MAD plus a small constant). The closer the consistency score is to 1, the more reliable the data.

[0127] Based on different but related data sources, cross-validation is performed using logical reasoning methods. For example, traffic flow data recorded by ETC gantry should be on the same order of magnitude as traffic flow data from nearby vehicle detectors; congestion events detected by video surveillance should correspond to an abnormal decrease in vehicle speed data for that road segment; rainfall data from meteorological displays should match the slippery conditions detected by road surface condition sensors. A rule base for logical relationships between heterogeneous data is established, including numerical correlation rules (e.g., the negative correlation between traffic flow and vehicle speed), event causality rules (e.g., rainfall causing a decrease in vehicle speed), and state consistency rules (e.g., multiple indicators in a congested area should be abnormal simultaneously). For each data point, relevant logical rules are searched based on its type, and other types of data that meet the rule conditions are searched to verify whether the logical relationship holds. A cross-validation score is calculated: 1 for a valid logical relationship, 0.5 for a partially valid relationship, and 0 for a invalid relationship.

[0128] Based on the results of consistency verification and cross-validation, a threshold-based method is used to identify conflicting data. A conflict identification threshold is set: data with a consistency score below 0.6 or a cross-validation score below 0.5 is marked as conflicting data. For data marked as conflicting, the conflict type (consistency conflict or logical conflict), conflict severity (score value), and information on other data sources involved in the conflict are recorded, constructing a conflict data record table. The frequency of conflicting data generated by each data source is statistically analyzed as a basis for subsequent reputation adjustments.

[0129] Based on data source reputation and conflict detection results, a weighted voting method is used to resolve data conflicts. For conflicting data groups (multiple data sources providing different observations on the same event or state), all participating data sources and their reputation scores are collected. A reputation-weighted voting mechanism is employed: the observation value of each data source is multiplied by its reputation score to obtain a weighted observation value; all weighted observation values ​​are summed and divided by the sum of all reputation scores to obtain a weighted average value as the resolved data value; for categorized data (such as event type judgment), the observation result from the data source with the highest reputation score is selected as the resolution result. The conflict resolution record is updated, including the original conflict data, resolution method, resolution result, and information on the data sources involved in the resolution.

[0130] Based on the results of validation and conflict resolution, the quality labeling information for each data point is updated. For data that passes consistency validation and cross-validation, its quality score is improved by averaging the original reasonableness score and the validation score as the new reasonableness score. For data that has undergone conflict resolution, its quality score is adjusted according to the confidence level of the resolution. If the resolution is based on a high-reputation data source, the quality score is maintained at a high level; if the resolution is based on a low-reputation data source, the quality score is appropriately reduced. For severely conflicting data that cannot be resolved, its quality score is set to 0 and it is marked as awaiting manual review. The reputation of data sources is updated, with negative adjustments made to data sources that generated conflicting data and positive adjustments made to data sources that provide consistent and reliable data.

[0131] The output of this module is a fused dataset that has undergone cross-validation and conflict resolution, with each data point containing updated quality labels and validation status information.

[0132] The results output module, based on the fused dataset, adopts a hierarchical classification output and quality feedback mechanism to obtain processing results for different application scenarios;

[0133] Based on the final dataset that has been validated and conflict resolved, a hierarchical classification output and quality feedback mechanism is adopted to obtain processing results for different application scenarios.

[0134] Based on the data priority classification, the final dataset is output to different processing queues according to priority. Level 1 (urgent) priority data is directly output to the real-time processing queue to trigger emergency response procedures, including automatic alarms, emergency dispatch, and on-site handling, with a processing latency requirement of less than 100 milliseconds. Level 2 (important) priority data is output to the priority processing queue for applications such as traffic status analysis, congestion warnings, and route optimization, with a processing latency requirement of less than 1 second. Level 3 (normal) priority data is output to the normal processing queue for applications such as statistical analysis, trend prediction, and report generation, with a processing latency requirement of less than 10 seconds. Level 4 (low priority) data is output to the batch processing queue for operations such as data cleaning, quality improvement, and historical archiving, with an acceptable processing latency of less than 1 minute.

[0135] Based on the data classification results and matching processing strategies, the data is distributed to the corresponding processing modules. Data for real-time processing strategies is sent to the streaming computing engine, using Apache Storm or Apache Flink for real-time stream processing. Data for time-series analysis strategies is sent to the time-series analysis module, where time-series forecasting algorithms are used for trend analysis and anomaly warning. Data for spatial analysis strategies is sent to the GIS analysis module for spatial calculations such as spatial interpolation, hotspot analysis, and path planning. Data for data cleaning strategies is sent to the data quality management module for quality improvement operations such as anomaly detection, missing value imputation, and noise filtering.

[0136] Collect the output results from each processing module, summarize the results, and standardize the format. For real-time processing results, generate event reports, early warning information, scheduling instructions, etc., and push them to relevant application systems via message queues or API interfaces. For analysis and processing results, generate analysis reports, prediction results, statistical charts, etc., and store them in the results database for subsequent querying and display. For data cleaning results, generate quality reports, repair records, cleaning statistics, etc., and update them to the data quality management system.

[0137] Based on processing results and user feedback, establish a continuous data quality improvement mechanism. Collect accuracy feedback on processing results, including indicators such as early warning hit rate, prediction accuracy, and anomaly detection accuracy. Analyze the performance of each data source under different application scenarios and update the parameters of the data source reputation assessment model. Analyze the effectiveness of processing strategies and optimize classification algorithms and strategy matching rules. Regularly generate data quality reports, including data source quality rankings, processing efficiency statistics, and system performance indicators, to provide a basis for system optimization.

[0138] The outputs of this module include processing results tailored to different application scenarios, such as real-time event response, traffic status analysis, and data quality reports.

[0139] A computer-readable storage medium for storing computer-readable instructions that, when read by a computer, enable the operation of a highway multi-source traffic data hierarchical classification and processing system.

[0140] In one embodiment of the present invention, a specific example is provided:

[0141] To better illustrate the technical solution of this invention, a 30-day field test was conducted on a highway. During the test, various data acquisition devices, including microwave vehicle detectors, high-definition video surveillance equipment, and meteorological monitoring stations, were deployed. The test area covered approximately 50 kilometers of road, encompassing various road scenarios such as service areas, toll stations, bridges, and tunnels. The test system processed over 1 million pieces of multi-source traffic data daily, verifying the effectiveness and practicality of the technical solution of this invention.

[0142] Data acquisition under normal traffic flow conditions:

[0143] At 9:00 AM on October 13, 2024, during the test period, the multi-source traffic data received by the system is shown in Table 1:

[0144] Table 1: Acquisition of multi-source data under normal traffic flow conditions;

[0145]

[0146] Data acquisition under abnormal traffic event conditions:

[0147] At 14:30 on October 15, 2024, during the testing period, the system detected a traffic anomaly at the K245 section. The received multi-source traffic data is shown in Table 2.

[0148] Table 2: Acquisition of multi-source data under abnormal traffic event conditions;

[0149]

[0150] Data processing under normal traffic flow conditions:

[0151] The system first converts the normal traffic flow data in Table 1 into a standardized format and queries the historical reputation scores of each data source. The reputation score for vehicle detector D_1523 is 0.85, the average reputation score for GPS data sources is 0.75, the reputation score for video surveillance V_0234 is 0.70, and the reputation score for weather station W_0089 is 0.90. Next, the system calculates the timeliness score for each data point. Using the current time 09:00:10 as a baseline, vehicle detector data with a 10-second delay has a timeliness score of e^(-0.01) multiplied by 10, which equals 0.905; GPS data with a delay of 5-17 seconds has a timeliness score of approximately 0.84-0.92; and weather data with a 10-minute delay has a timeliness score of e^(-0.001) multiplied by 600, which equals 0.549. During the data quality check, all required fields of the data were complete, and the completeness score was 1.0. The historical vehicle speed range of this road section was 60-120 km / h, and all observations were within a reasonable range, with a reasonableness score of 1.0.

[0152] The system then calculates the traffic urgency level of section K238. The current traffic flow of 1200 vehicles / h is lower than the historical average of 1150 vehicles / h, resulting in a traffic flow anomaly of (1200 - 1150) divided by (1150 + 1), which equals 0.043. The current speed of 85 km / h is within the normal range of [70, 110] km / h, so the speed anomaly is 0. The current visibility of 8000m is good, and the weather risk score is 0.4 multiplied by (1 - 8000 divided by 10000) plus 0 plus 0, which equals 0.08. There are no emergency event keywords, so the emergency event score is 0. The overall urgency level is 0 multiplied by 0.5 plus (0.043 plus 0 plus 0.08) divided by 3 multiplied by 0.5, which equals 0.021. This section is in a normal state with a low urgency level.

[0153] During the priority assessment phase, the system calculates a comprehensive priority score for each data item. Taking vehicle speed data from a vehicle detector as an example, the comprehensive priority score is 0.2 multiplied by 0.85 plus 0.2 multiplied by 0.905 plus 0.2 multiplied by (1.0 multiplied by 1.0) plus 0.4 multiplied by 0.021, which equals 0.559, classifying it as Level 3 (regular) priority. The priority score for GPS vehicle speed data is approximately 0.52-0.56, also at Level 3 priority. Traffic flow data and meteorological data have similar priority scores, both at Level 3.

[0154] The system performs adaptive classification based on data characteristics. Vehicle speed data has spatial continuity characteristics (vehicle speeds are similar at adjacent locations) and is classified as spatial continuous data, with a spatial interpolation processing strategy configured. Traffic flow data has temporal periodic characteristics and is classified as periodic data, with a time series analysis strategy configured.

[0155] In the multi-source data fusion stage, the system identifies four vehicle speed data points describing the same road segment's speed parameters, establishing a data association set {85, 82, 88, 78} km / h. First, conflict detection is performed, with a standard deviation of 4.11 km / h, less than the threshold of 10 km / h, indicating the data is essentially consistent with no serious conflicts. Then, the fusion weights (reputation × timeliness) of each data source are calculated: vehicle detector weight 0.769, GPS1 weight 0.690, GPS2 weight 0.630, and video weight 0.623. Next, the weighted fused vehicle speed is calculated: fused vehicle speed = (85 × 0.769 + 82 × 0.690 + 88 × 0.630 + 78 × 0.623) ÷ (0.769 + 0.690 + 0.630 + 0.623) = 83.5 km / h. Finally, the fused vehicle speed for the current time segment is output as 83.5 km / h, with a confidence level of 0.9.

[0156] Data processing procedures under abnormal traffic event conditions:

[0157] The system first converts the abnormal traffic event data in Table 2 into a standardized format and queries the historical reputation scores of each data source. The reputation score for vehicle detector D_1567 is 0.88, the average reputation score for GPS data sources is 0.72, the reputation score for video surveillance V_0245 is 0.85, and the reputation score for weather station W_0156 is 0.92. Next, the system calculates the timeliness score for each data point. Using the current time 14:30:00 as the baseline, the vehicle detector data with a 5-second delay has a timeliness score of e^(-0.01)^5, which equals 0.951; GPS data with an 8-15 second delay has a timeliness score of approximately 0.86-0.92; video data with a 3-second delay has a timeliness score of 0.970; and weather data with a 5-minute delay has a timeliness score of e^(-0.001)^300, which equals 0.741. During the data quality check, all required fields of the data were complete, and the completeness score was 1.0. The historical speed range of this road section was 60-120 km / h, and the current observed value of 15-25 km / h was obviously abnormal, but it was consistent with the characteristics of an accident scenario, and the reasonableness score was 0.8.

[0158] The system then calculates the traffic urgency level of section K245. The current traffic flow of 300 vehicles / h is significantly lower than the historical average of 1100 vehicles / h, with an anomaly of (300 - 1100) divided by (1100 + 1) equaling 0.727. The current speed of 20 km / h is far below the normal range of [70, 110] km / h, with an anomaly of (70 - 20) divided by 70 equaling 0.714. The current visibility of 3000m is poor, and the weather risk score is 0.4 multiplied by (1 - 3000 divided by 10000) plus 0.3 multiplied by 1 plus 0 equaling 0.58. The keyword "accident" is detected, and the emergency event score is 0.9. The overall urgency level is 0.9 multiplied by 0.5 plus (0.727 + 0.714 + 0.58) divided by 3 multiplied by 0.5 equaling 0.787, indicating that this section is in a highly urgent state.

[0159] During the priority assessment phase, the system calculates a comprehensive priority score for each data item. Taking vehicle speed data from a vehicle detector as an example, the comprehensive priority score is 0.2 x 0.88 + 0.2 x 0.951 + 0.2 x (1.0 x 0.8) + 0.4 x 0.787 = 0.841, which is classified as Level 1 (urgent) priority. The priority score for video surveillance data is 0.2 x 0.85 + 0.2 x 0.970 + 0.2 x (1.0 x 1.0) + 0.4 x 0.787 = 0.879, also Level 1 priority. The priority scores for GPS data and meteorological data are approximately 0.75-0.82, classifying them as Level 2 (important) priority.

[0160] The system performs adaptive classification based on data characteristics. Vehicle speed data has abrupt change anomaly characteristics (sharp drop in vehicle speed) and is classified as abrupt change data category, with anomaly detection and rapid response processing strategies configured. Traffic flow data has time-series anomaly characteristics and is classified as abrupt time-series data category, with anomaly pattern recognition strategies configured.

[0161] In the multi-source data fusion stage, the system identified four abnormal vehicle speed parameters describing the same road segment, establishing a data association set {15, 20, 25, 18} km / h. First, conflict detection was performed, with a standard deviation of 4.24 km / h, less than the threshold of 10 km / h, indicating that the data were basically consistent and had no serious conflicts. Then, the fusion weights of each data source (reputation × timeliness × urgency weight) were calculated: vehicle detector weight 0.837, video weight 0.825, GPS1 weight 0.621, and GPS2 weight 0.590. Next, the weighted fused vehicle speed was calculated: fused vehicle speed = (15 × 0.837 + 20 × 0.825 + 25 × 0.621 + 18 × 0.590) ÷ (0.837 + 0.825 + 0.621 + 0.590) = 19.2 km / h. The final output shows that the current vehicle speed on this road segment is 19.2 km / h, with a confidence level of 0.95, and the accident warning mechanism is immediately triggered.

[0162] By comparing the processing procedures under two different traffic conditions, it can be seen that the system can automatically adjust its processing strategy according to different traffic scenarios: under normal traffic flow scenarios, the system prioritizes data quality and timeliness, adopting a conventional processing flow; under abnormal traffic event scenarios, the system significantly increases the urgency weight, prioritizes processing high-reliability data sources, and triggers a rapid response mechanism. This adaptive processing capability ensures that the system can provide accurate and timely data support under various traffic conditions.

[0163] This invention achieves intelligent and refined processing of multi-source traffic data on highways by establishing a four-dimensional dynamic hierarchical mechanism, an adaptive data classification method, and an evidence-based fusion framework. The system can dynamically adjust processing priorities based on data source reliability, data quality, timeliness, and the urgency of traffic scenarios, ensuring millisecond-level response times for emergency events. Automatic data feature identification and optimal processing strategy matching improve the accuracy and efficiency of data processing. Multi-source cross-validation and conflict resolution effectively enhance the reliability of fused data. This invention provides an effective technical means for intelligent highway management and has significant application value.

[0164] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.

Claims

1. A highway multi-source traffic data hierarchical classification processing system, characterized in that, The application relates to a traffic data processing method and device. The method comprises the following steps: a data acquisition standardization module is used to acquire multi-source traffic data, pre-process the multi-source traffic data, and obtain a standardized data set; a data source quality evaluation module is used to evaluate the credibility of each data source based on the standardized data set, the credibility evaluation is based on the historical data accuracy, data reporting stability and data integrity of the data source, a weighted linear combination method is used to calculate the comprehensive credibility of the data source, the data source credibility is dynamically updated according to the data verification result, a positive incentive updating mechanism is used to improve the credibility when the data passes the verification, and a negative punishment updating mechanism is used to reduce the credibility when the data fails the verification; the time effectiveness, integrity and rationality of each piece of data are scored, the time effectiveness score is calculated by using an exponential decay function, the integrity score is calculated by using a weighted field integrity method, and the rationality score is calculated by using a segmented scoring rule based on historical statistical range testing, and an enhanced data set is obtained; a scene emergency degree calculation module is used to calculate the traffic anomaly degree, the vehicle speed anomaly degree and the weather risk score based on the enhanced data set, and to identify an emergency event, so as to obtain a traffic scene emergency degree score; a data grading module is used to calculate the comprehensive priority score of the data based on the enhanced data set and the traffic scene emergency degree score, the comprehensive priority score is based on four dimensions of data source credibility, data time effectiveness, data quality and traffic scene emergency degree, a weighted linear combination method is used for calculation, wherein the integrity score and the rationality score of the data are multiplied to obtain a comprehensive data quality score, and a threshold segmentation method is used to obtain a graded data set; a data classification module is used to extract multi-dimensional features of the data based on the graded data set, and to perform adaptive classification, so as to obtain a classified data set; a data fusion module is used to identify multi-source data of the same traffic parameter based on the classified data set and to perform conflict detection, to process the conflict data by using evidence theory fusion, and to obtain a fusion data set; 2.The highway multi-source traffic data hierarchical classification processing system of claim 1, wherein, a result output module is used to obtain processing results for different application scenarios by using a hierarchical classification output and quality feedback mechanism based on the fusion data set. The calculation of the traffic scene emergency degree score comprises the following steps: a deviation degree calculation method is used to evaluate the traffic anomaly degree based on the traffic flow data of the current road section; a segmented deviation degree calculation method is used to evaluate the vehicle speed anomaly degree based on the vehicle speed data of the current road section, the anomaly degree is a first value when the vehicle speed is within a preset normal range, and the anomaly degree is a relative deviation ratio when the vehicle speed exceeds the normal range; a multi-factor weighting method is used to calculate the weather risk score based on the weather monitoring data; an emergency event is identified by using a keyword matching method based on the data text description field, and a keyword library containing preset emergency event keywords is constructed; 3.The highway multi-source traffic data hierarchical classification processing system of claim 1, wherein, a hierarchical weighting method is used to calculate the comprehensive emergency degree of the traffic scene, the emergency degree is not lower than a first threshold value when there is an emergency event, and the emergency degree reaches a second threshold value when there is no emergency event but the conventional indicators are severely abnormal. The adaptive classification comprises the following steps: statistical features of the data are extracted by using a statistical analysis method based on the numerical fields of the data; time sequence features are extracted by using a frequency domain analysis method based on the time sequence attributes of the data; Based on the data geographical location information, the spatial statistical analysis method is used to extract spatial characteristics; The statistical characteristics, time series characteristics and spatial characteristics are combined into a comprehensive feature vector; The clustering algorithm is used to classify the comprehensive feature vector to obtain a data set with classification identification.

4. The highway multi-source traffic data hierarchical classification processing system of claim 1, wherein, The conflict data is fused and processed by using the evidence theory, which includes: Based on the time stamp, geographical location and data parameter type, the multi-source data describing the same traffic parameter is identified and classified into a data association set; The consistency of the multiple data in the data association set is tested to determine whether there is a conflict; For the data with conflicts, the weighted voting method is used to eliminate the data conflicts, and the weighted average value is calculated based on the credibility of each data source; The final fusion result is obtained by weighted fusion calculation.

5. The highway multi-source traffic data hierarchical classification processing system of claim 1, wherein, The threshold segmentation method is used to obtain the hierarchical data set, which includes: A plurality of priority threshold values are set to divide the comprehensive priority score interval into different levels; According to the comparison result of the comprehensive priority score of the data and the threshold value, the corresponding priority label is assigned to each data; The priority level identification field and the priority score field are added to each data to obtain a data set with hierarchical identification.

6. The highway multi-source traffic data hierarchical classification processing system of claim 1, wherein, The calculation of the comprehensive priority score includes: Based on the historical accuracy, stability and integrity of the data source, the weighted linear combination method is used to calculate the data source credibility; Based on the time difference between the data collection time and the current time, the exponential decay function is used to calculate the data timeliness score; Based on the data field integrity and numerical rationality, the multiplication combination method is used to calculate the data quality score; The data source credibility, data timeliness score, data quality score and traffic scene emergency degree are weighted and summed according to the set weight to obtain the data comprehensive priority score.

7. A computer-readable storage medium, characterized in that, It is used to store computer readable instructions, when the computer readable instructions are read by the computer, it can run a highway multi-source traffic data hierarchical classification processing system as claimed in any one of claims 1-6.

Citation Information

Patent Citations

  • Highway event handling method based on big data

    CN120510721A