Method and system for discovering traffic abnormal data based on big data real-time analysis
Through the real-time analysis method of big data, Flume and Flink are used to build a data hierarchical architecture, which solves the lag problem of abnormal data detection in the highway gantry system, realizes real-time detection and efficient processing of highway traffic abnormal data, and improves the accuracy of abnormal detection and the intelligence level of the system.
Patent Information
- Application Number
- CN202510704393.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art cannot effectively discover abnormalities in vehicle traffic paths in highway gantry systems in real time, resulting in data processing lag, low analysis efficiency, and failure to detect abnormalities in time, increasing system pressure and leading to toll losses.
Using a real-time analysis method based on big data, data is collected through Flume and synchronized to Kafka, and Flink is used to build a data hierarchical architecture for intelligent hierarchical analysis, including ODS, DIM, DWD and DWS layers, data cleaning, association and abnormal analysis are carried out, multi-dimensional abnormal data sets are generated and real-time query and early warning are performed.
Real-time detection of highway traffic abnormal data is realized, the accuracy and response speed of abnormal detection is improved, toll losses are reduced, customer complaints are reduced, data processing efficiency and system intelligence are improved.
Smart Images

Figure CN120388472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent transportation big data, and particularly to a method and system for discovering traffic abnormal data based on real-time big data analysis. Background Art
[0002] Since the removal of provincial boundary toll stations on national expressways, expressway gantry systems have been vigorously constructed across the country. The national expressways have entered a new mode of charging through gantries, marking vehicle passing trajectories, and accumulating gantry amounts at the exit for toll collection based on ETC or CPC. This mode has broken the past traditional toll collection mode of point-to-point at the entrance and exit stations of each province, making the vehicle passing routes clearer and the toll collection more accurate.
[0003] During the operation of the gantry system, a large amount of transaction and license plate recognition data is generated. However, there are many problems with these data. At the data level, affected by the environment and the system, there are situations such as incorrect license plate recognition, incorrect time synchronization, and mis-triggered abnormal shooting in the gantry license plate recognition data. At the same time, there are also problems of data upload delay and missed transaction transmission, and it is difficult to accurately restore the route in the massive data. In terms of processing methods, traditional methods rely on post-event reconciliation and auditing, unable to detect abnormalities in real time, resulting in processing delays. In terms of performance, the massive data poses great challenges to analysis, not only with low analysis efficiency but also causing huge pressure on subsequent systems. In the case where the daily data volume reaches millions, even tens of millions, or hundreds of millions, how to accurately locate abnormalities from these data, reduce toll losses, reduce customer complaints, and reasonably utilize license plate recognition data without being interfered by incorrect data has always been the focus and difficulty of the work in each province. The prior art has not effectively solved the above problems, so there is an urgent need for an efficient and real-time abnormal data detection solution. Summary of the Invention
[0004] In view of the problems of lagging data processing and low efficiency of abnormal location in the prior art, the present invention provides a method for discovering traffic abnormal data based on real-time big data analysis. Through big data methods, real-time analysis is carried out on traffic data such as gantries, license plate recognition, and entrance and exit transactions generated on expressways, quickly and efficiently discovering abnormal data, warning of suspected abnormal vehicle passing data, and providing efficient data support for reconciliation and auditing. The present invention also relates to a system for discovering traffic abnormal data based on real-time big data analysis.
[0005] The technical solution of the present invention is as follows:
[0006] A method for discovering traffic abnormal data based on real-time big data analysis, characterized by comprising the following steps:
[0007] Steps for multi-dimensional traffic data collection: Real-time collect a large amount of gantry transaction data, entrance and exit transaction data, and gantry license plate recognition data generated by highways through Flume, and synchronize the big data collected in real time to different corresponding Topics in Kafka according to the data type or collection source;
[0008] Steps for intelligent hierarchical analysis of traffic data: Build a data hierarchical architecture based on Flink for intelligent hierarchical analysis of traffic data, including the ODS layer, DIM layer, DWD layer, and DWS layer in sequence. Among them, the ODS layer: stores the original big data collected in real time; the DIM layer: constructs several dimension tables including gantry association information table, gantry basic information table, holiday information table, and vehicle type information table, and pre-loads the corresponding dimension tables to the Redis cache according to scheduled tasks or trigger instructions; the DWD layer: cleans the original big data collected in real time, and associates the cleaned data with the corresponding dimension tables in the DIM layer through specified fields to generate standardized traffic data; the DWS layer processes data in three stages:
[0009] a. Fusion stage: Read the standardized traffic data generated by the DWD layer from Kafka, associate it according to the license plate, cut the multiple passages of the same vehicle into multiple passage data according to the time sequence and traffic logic, and then publish the processed multiple passage data to Kafka;
[0010] b. Analysis stage: Read the multiple passage data published in the fusion stage from Kafka, and based on the gantry association information table and gantry basic information table in the DIM layer, perform real-time analysis of abnormal indicators including no exit at the entrance, inconsistent cumulative gantry amount and exit amount, abnormal path, and license plate recognition error through the Flink real-time data stream processing engine, and publish the analysis results to Kafka and store them in the Doris database at the same time; During the real-time analysis process, based on the holiday information in the holiday information table in the DIM layer and the vehicle type data information in the vehicle type information table, automatically identify whether it is a holiday free vehicle or free flow. If it is abnormal data generated by a holiday free vehicle, it is automatically marked as a special abnormality. If it is free flow, the flow is automatically determined to be normal;
[0011] c. Statistics stage: Based on the Flink node, receive the analysis results pushed in the analysis stage from Kafka, perform aggregation statistics according to the preset dimensions, and generate a vehicle abnormal detail data set and abnormal information, various dimension statistical information, a detail data set and analysis results after aggregation of a single vehicle passage;
[0012] Steps for multi-dimensional output of traffic abnormal data: Provide the data generated in the statistics stage through the API interface, supporting multi-dimensional real-time query and warning.
[0013] Preferably, in the step of intelligent hierarchical analysis of traffic data, the gantry association information table constructed by the DIM layer includes several field combinations of gantry number, superior gantry number, inferior gantry number, section identification, and geographical coordinates. The gantry basic information table includes several field combinations of gantry number, gantry name, toll type, billing rule version, and equipment status. The holiday information table includes several field combinations of holiday date range, holiday type, free vehicle type category, and free policy description. The vehicle type information table includes several field combinations of vehicle type code, vehicle type name, and toll coefficient.
[0014] Preferably, in the step of intelligent hierarchical analysis of traffic data, the DWD layer cleans the originally collected big data in real time, including uniformly converting the timestamps of different collection sources into Beijing time format through the Flink parallel computing engine, filtering records with negative or transaction amounts exceeding a preset threshold, completing or discarding records missing the gantry number, and judging and removing duplicate records based on the transaction serial number or timestamp; associating the cleaned data with the corresponding dimension tables in the DIM layer through specified fields, specifically: associating the gantry number field in the cleaned data with the gantry number fields in the gantry association information table and the gantry basic information table in the DIM layer, associating the transaction time field in the cleaned data with the holiday date range field in the holiday information table in the DIM layer, and associating the vehicle type code field in the cleaned data with the vehicle type code field in the vehicle type information table in the DIM layer.
[0015] Preferably, in the analysis stage of the DWS layer, real-time analysis of abnormal indicators such as outlier data, gantry time jumps, having license plate recognition but no transaction, having virtual gantry data but no entrance / exit transaction data, OBU or CPC cumulative amount billing but no gantry transaction data, license plate recognition reverse marking, and license plate recognition mislabeling is also performed through the Flink real-time data stream processing engine.
[0016] Preferably, in the statistical stage of the DWS layer, aggregation statistics are performed according to preset dimensions. The preset dimensions include abnormal type, time window, and toll station, where the time window includes time windows in units of hours / days / weeks; the generated vehicle abnormal detail data set and abnormal information include the time, location, type, and vehicle identification involved in the abnormality. The statistical information for each dimension includes the statistical results aggregated by time window, abnormal type, and toll station dimensions. The detail data set and analysis results after aggregation of a single vehicle passage include the complete passage path and billing details.
[0017] Preferably, in the step of multi-dimensional output of traffic anomaly data, the vehicle anomaly detail data set, anomaly information, statistical information of each dimension generated in the statistical stage, as well as the detail data set and analysis results after aggregation of single vehicle passes are provided externally through an API interface. The API interface adopts load balancing technology and data caching strategy, supports HTTP / HTTPS / RPC communication protocols, and supports two data interaction modes of active push and passive pull, and also supports multi-dimensional real-time query and warning.
[0018] A system for discovering traffic anomaly data based on real-time big data analysis, characterized by comprising a traffic multi-dimensional data collection module, a traffic data intelligent hierarchical parsing module, and a traffic anomaly data multi-dimensional output module connected in sequence.
[0019] The traffic multi-dimensional data collection module: Real-time collects a large amount of gantry transaction data, entrance and exit transaction data, and gantry license plate recognition data generated on highways through Flume, and synchronizes the real-time collected big data to different Topics in Kafka respectively according to the data type or collection source.
[0020] The traffic data intelligent hierarchical parsing module: Constructs a data hierarchical architecture based on Flink for intelligent hierarchical parsing of traffic data, including an ODS layer, a DIM layer, a DWD layer, and a DWS layer in sequence. Among them, the ODS layer: Stores the original big data collected in real time; the DIM layer: Constructs several dimension tables including a gantry association information table, a gantry basic information table, a holiday information table, and a vehicle type information table, and pre-loads the corresponding dimension tables into the Redis cache according to a scheduled task or a trigger instruction; the DWD layer: Cleans the original big data collected in real time, and associates the cleaned data with the corresponding dimension tables in the DIM layer through specified fields to generate standardized traffic data; the DWS layer processes the data in three stages:
[0021] a. Fusion stage: Reads the standardized traffic data generated by the DWD layer from Kafka, associates them according to the license plate, cuts the multiple passes of the same vehicle into multiple segments of traffic data in chronological order and traffic logic, and then publishes the processed multiple segments of traffic data to Kafka.
[0022] b. Analysis phase: Read multiple segments of passing data published in the fusion phase from Kafka. Based on the gantry association information table and gantry basic information table in the DIM layer, use the Flink real-time data stream processing engine to perform real-time analysis of abnormal indicators including entry without exit, inconsistent gantry cumulative amount and exit amount, path anomaly, and license plate recognition error, and publish the analysis results to Kafka, while storing them in the Doris database; during the real-time analysis process, based on the holiday information in the holiday information table in the DIM layer and the vehicle type data information in the vehicle type information table, automatically identify whether it is a holiday free vehicle type or free flow. If the abnormal data is generated by a holiday free vehicle type, it is automatically marked as a special anomaly. If it is free flow, the flow is automatically determined to be normal;
[0023] c. Statistics phase: Based on the Flink node, receive the analysis results pushed by the analysis phase from Kafka, perform aggregation statistics according to preset dimensions, and generate a vehicle anomaly detail data set and anomaly information, statistical information for each dimension, a detail data set and analysis results after aggregating a single vehicle passage;
[0024] The traffic anomaly data multi-dimensional output module: Provide the data generated in the statistics phase externally through the API interface, supporting multi-dimensional real-time query and warning.
[0025] Preferably, in the traffic data intelligent hierarchical parsing module, the gantry association information table constructed in the DIM layer includes several field combinations such as gantry number, superior gantry number, inferior gantry number, section identification, and geographical coordinates. The gantry basic information table includes several field combinations such as gantry number, gantry name, charging type, billing rule version, and equipment status. The holiday information table includes several field combinations such as holiday date range, holiday type, free vehicle type category, and free policy description. The vehicle type information table includes several field combinations such as vehicle type code, vehicle type name, and charging coefficient;
[0026] The DWD layer cleans the originally collected big data, including uniformly converting the timestamps from different collection sources to the Beijing time format through the Flink parallel computing engine, filtering records with negative or transaction amounts exceeding the preset threshold, completing or discarding records with missing gantry numbers, and judging and removing duplicate records based on the transaction serial number or timestamp; associate the cleaned data with the corresponding dimension tables in the DIM layer through specified fields. Specifically: associate the gantry number field in the cleaned data with the gantry number fields in the gantry association information table and gantry basic information table in the DIM layer, associate the transaction time field in the cleaned data with the holiday date range field in the holiday information table in the DIM layer, and associate the vehicle type code field in the cleaned data with the vehicle type code field in the vehicle type information table in the DIM layer.
[0027] Preferably, in the traffic data intelligent hierarchical parsing module, during the analysis stage of the DWS layer, the Flink real-time data stream processing engine is also used to perform real-time analysis of abnormal indicators such as outlier data, gantry time jumps, license plate recognition without transactions, virtual gantry data without entrance / exit transaction data, OBU or CPC cumulative amount billing without gantry transaction data, license plate recognition reverse labeling, and license plate recognition mislabeling.
[0028] Preferably, in the traffic data intelligent hierarchical parsing module, during the statistical stage of the DWS layer, aggregation statistics are performed according to preset dimensions. The preset dimensions include abnormal types, time windows, and toll stations, where the time windows include time windows in units of hours / days / weeks; the generated vehicle abnormal detail data set and abnormal information include the time, location, type, and vehicle identification involved in the abnormality. The statistical information for each dimension includes the statistical results aggregated by time window, abnormal type, and toll station dimensions. The detailed data set and analysis results after aggregation of a single vehicle passage include the complete passage path and billing details.
[0029] The beneficial effects of the present invention are:
[0030] The present invention provides a method for discovering traffic abnormal data based on real-time big data analysis. In the traffic multi-dimensional data collection step, a distributed data collection tool Flume is used to achieve millisecond-level data collection, ensuring the real-time nature of a large amount of traffic data (such as gantry transaction records, entrance and exit transaction data, license plate recognition information), meeting the requirements for rapid response to abnormal events on expressways. Using the Topic partitioning mechanism of Kafka message middleware, different types of data (such as transaction data, license plate recognition data) are isolated and stored distributively, supporting parallel processing of TB-level data, and being able to achieve asynchronous data processing with higher throughput and efficiency. Classified storage is based on data types or collection sources, facilitating seamless expansion when new data sources (such as ETC transaction data) are added subsequently.The intelligent hierarchical parsing steps of traffic data are based on Flink to build a data hierarchical architecture for intelligent hierarchical parsing of traffic data. This hierarchical architecture can make the data system clearer and simplify complex problems. At the same time, stream processing can process data in real time and incrementally, and aggregation analysis can statistically analyze data more quickly and conveniently. The ODS layer stores all raw data to achieve the integrity of raw data, supports data traceability and historical rollback, avoids the risk of data loss, and can also store PB-level data through distributed file systems such as HDFS to achieve distributed storage and ensure high availability. The DIM layer constructs several dimension tables including gantry association information tables, gantry basic information tables, holiday information tables, and vehicle type information tables. Since the data volume of this type of data is small and the changes are minor, when the service starts, the full amount of data is directly imported into the redis cache. When there are changes, the cache is directly updated. The dimension tables are pre-loaded into the Redis cache, reducing the data association query latency from seconds to milliseconds, improving the real-time analysis efficiency of the solution. At the same time, the holiday information table and vehicle type information table provide unified business rules for anomaly analysis, realizing the reuse of business rules and avoiding repeated calculations. The DWD layer converts heterogeneous raw data into a unified format (such as timestamp standardization and vehicle type code unification) through cleaning and association operations, eliminates data ambiguity, realizes data standardization, and can improve the analysis accuracy of the solution. The DWS layer processes data in three stages. In its fusion stage, based on license plate association of multi-source data, it fuses gantry transaction data, license plate recognition data, and gantry association information tables, sorts the multiple passing records of the same vehicle in chronological order, and constructs a complete vehicle passing trajectory in real time from massive data, providing a basis for anomaly analysis, cutting and structuring multiple segments of passing data, reducing redundant information, and reducing the subsequent processing pressure. In the analysis stage, the Flink real-time data stream processing engine can be used to achieve millisecond-level anomaly recognition, support real-time warning, and based on the Kafka message queue mechanism and Flink's state management, detect and compensate for data upload delays and missed transactions. In its analysis stage, by analyzing the license plate recognition error metrics in real time, it automatically identifies and marks data such as abnormal trigger misfires and license plate recognition errors, combines the vehicle type information table and holiday information table in the DIM layer, corrects or filters the error data, improves the quality of raw data, thus solving problems such as license plate recognition errors, time alignment errors, and abnormal trigger misfires in the existing gantry license plate recognition data, automatically identifying free vehicle types and flows during holidays, reducing manual intervention, and reducing the misjudgment rate. Its statistical stage supports multi-dimensional aggregation by time (such as hour / day / week), space (such as road section / toll station), business type (such as anomaly type), etc., meets different decision-making needs, generates vehicle anomaly detail data sets and anomaly information, statistical information for each dimension, detail data sets and analysis results after aggregation of single vehicle passes, and at the same time outputs the original anomaly details and aggregation statistical results, taking into account both refined investigation and macro decision-making.The multi-dimensional output step of its traffic anomaly data enables multi-channel access, multi-dimensional real-time query and warning, real-time interaction, meets the needs of emergency event handling, and provides visualization support, assisting managers in quickly locating problems and enhancing the decision-making support ability of the solution.
[0031] The present invention provides a method for discovering traffic anomaly data based on real-time big data analysis. By introducing real-time big data warehouse technology, it integrates gantry transaction data, entrance / exit transaction data, and gantry license plate recognition data for analysis, and can more efficiently and quickly discover situations such as data upload delay, missing transmission, abnormal transaction amount, and abnormal passing trajectory, thus saving losses for users and road section operation units, and also providing data support for subsequent auditing and reconciliation services. At the technical level, doris (a high-performance real-time analysis database with MPP architecture) is introduced to store and query data more quickly; the flink distributed streaming computing framework is introduced to analyze data more real-time; the redis bypass cache technology is introduced to solve the performance problems caused by frequent database queries during the analysis process. Due to situations such as users shielding ETC cards, CPC cards, damaged cards, and abnormal passing in actual passing, there is no gantry transaction data, which brings various difficulties to the traditional method of restoring vehicle paths only using gantry transaction data. Therefore, at the business level, gantry license plate recognition data is introduced and combined with gantry transaction data to more accurately restore vehicle passing paths. At the same time, the data of the same passing is fused and stored, and the entrance / exit transaction data, gantry transaction data, and gantry license plate recognition data of the same passing can be retrieved in one query, significantly improving the data query efficiency.
[0032] The present invention realizes real-time analysis of traffic data such as gantries, license plate recognitions, and entrance / exit flows generated on highways through big data methods, and can classify and type-warning suspected abnormal vehicle passing data. It has full-link real-time performance: the full-process delay from data collection to anomaly output is controlled within seconds. Compared with the traditional T+1 analysis mode, the response speed is increased by several orders of magnitude, supporting the real-time discovery and handling of highway anomaly events (such as toll evasion, equipment failure); data value mining: through a hierarchical architecture and multi-dimensional analysis, a large amount of raw data is transformed into decision-making information (such as anomaly trends, road section risk assessment). The aggregated data of a single passing supports derivative applications such as path optimization and toll auditing, enhancing the application value of the solution; scalability and flexibility: the distributed architecture (Flume + Kafka + Flink) supports horizontal expansion and can easily handle the sharp increase in data volume brought by the growth of traffic flow; business rule automation: automatically recognizes holiday free policies and vehicle types, reduces manual intervention, avoids human errors, improves toll collection accuracy, and the anomaly marking and classification mechanism (such as special anomalies, normal flows) provides clear guidance for subsequent processing.
[0033] Furthermore, by clarifying the field combinations of each dimension table, a unified data model is provided for subsequent data association and analysis, eliminating semantic differences between different data sources. The structured field design supports efficient JOIN operations. For example, through the gantry number field, the gantry association information table and the gantry basic information table can be quickly associated, improving the association efficiency and reducing the time complexity of data processing. Business rules such as toll policies (holiday information table) and vehicle type classifications (vehicle type information table) are stored in data form, thus solidifying the business rules, facilitating the automatic execution of decision-making logic, and reducing manual intervention.
[0034] Furthermore, through timestamp standardization, outlier filtering, missing value handling, and deduplication operations, the accuracy and integrity of the original data are significantly improved, providing reliable data quality assurance for subsequent analysis. The association relationships of specified fields (such as gantry number, transaction time, vehicle type code) are clearly defined, ensuring the accuracy of association and enabling the cleaned data to be precisely matched with the dimension table, avoiding data association ambiguity. Cleaning and association are completed before the data enters the analysis stage, reducing the processing volume of invalid data, lowering the resource consumption of subsequent computing nodes, and optimizing computing resources.
[0035] Furthermore, real-time analysis of multiple abnormal indicators such as isolated point data and gantry time jumps is added to ensure the comprehensiveness of abnormal detection, covering more traffic abnormal scenarios, enhancing the system's risk identification ability, and enabling the detection of special abnormal situations in emerging technology applications such as ETC / OBC devices and virtual gantries, adapting to the technological evolution of the highway toll system. Through the Flink real-time stream processing engine, millisecond-level abnormal detection is achieved, improving the warning timeliness. Compared with traditional batch detection methods, potential problems can be discovered hours or even days in advance.
[0036] Furthermore, multi-dimensional aggregation statistics such as abnormal types, time windows, toll stations, etc. are provided to meet the decision-making needs of different management levels. For example, analyze the abnormal trend by hour and locate the high-incidence areas of problems by toll station. Statistical results are automatically generated through predefined time windows (hours / days / weeks), and periodic business insights can be obtained without manual intervention, improving the data analysis efficiency. The complete travel path and billing details are output, providing direct evidence for business scenarios such as toll auditing and customer complaint handling, and supporting the closed-loop management of business processes.
[0037] Furthermore, the API interface adopts load balancing technology and data caching strategy. The load balancing technology ensures the stability of the API service under high concurrency, and the data caching strategy reduces repeated calculations and improves the response speed. The system availability can reach 99.99%; it supports multiple communication protocols such as HTTP / HTTPS / RPC, is compatible with the Web side, mobile side and third-party systems, and expands the application scope of data services; at the same time, it has interactive flexibility, supports dual modes of active push (such as exception warning) and passive pull (such as historical query), meets the different scenario requirements of real-time monitoring and post-event analysis, and improves the user experience.
[0038] The present invention also relates to a system for discovering traffic abnormal data based on big data real-time analysis. This system corresponds to the above-mentioned method for discovering traffic abnormal data based on big data real-time analysis, and can be understood as a system for implementing the method for discovering traffic abnormal data based on big data real-time analysis. It includes a traffic multi-dimensional data collection module, a traffic data intelligent hierarchical parsing module, and a traffic abnormal data multi-dimensional output module that are connected in sequence. Through the traffic multi-dimensional data collection module, a large amount of traffic multi-dimensional data is collected in real time. The traffic data intelligent hierarchical parsing module cleans, correlates, and analyzes the data to accurately identify various abnormal indicators. The traffic abnormal data multi-dimensional output module quickly outputs abnormal details and statistical results, and supports multi-dimensional real-time query and warning. The present invention realizes millisecond-level response through stream processing technology, has the advantage of real-time performance, and solves the problem that traditional solutions rely on batch processing (such as daily timing analysis) and cannot meet the real-time monitoring requirements of expressways; the present invention integrates multi-source data such as a large amount of gantry transaction data, entrance and exit transaction data, and gantry license plate recognition data, has data processing depth, and solves the problem that traditional solutions only analyze a single data source (such as transaction data or license plate recognition data), and the abnormal detection accuracy is increased by more than 50%; the present invention automatically generates analysis results through predefined dimensions through intelligent hierarchical parsing, and supports multi-dimensional query, which improves the decision-making assistance ability of the solution and solves the problem that traditional solutions only output raw data or simple statistics. Each module closely cooperates to realize the full-process automation and real-timeization from data collection to abnormal processing. Compared with traditional methods, the response speed is increased by several orders of magnitude, the abnormal detection accuracy is increased by more than 50%, the toll loss is effectively reduced, the customer complaint rate is reduced, and the operation management efficiency and intelligent level of expressways are significantly improved. Brief Description of the Drawings
[0039] Figure 1 is a flowchart of the method for discovering traffic abnormal data based on big data real-time analysis of the present invention.
[0040] Figure 2 is a working principle diagram of the Kafka message queue of the present invention.
[0041] Figure 3It is the data hierarchical architecture diagram constructed in the intelligent hierarchical parsing steps of traffic data in the present invention.
[0042] Figure 4 It is the working principle diagram of Flink in the present invention.
[0043] Figure 5 It is the preferred flow chart of the method for discovering traffic abnormal data based on real-time big data analysis in the present invention.
[0044] Figure 6 It is the flow chart of the intelligent hierarchical parsing steps of traffic data based on Flink in the present invention.
[0045] Figure 7 It is the preferred flow chart of the analysis stage of the DWS layer based on Flink in the present invention.
[0046] Figure 8 It is the overall preferred architecture diagram of the system for discovering traffic abnormal data based on real-time big data analysis in the present invention. Detailed implementation manners
[0047] The present invention will be described below with reference to the accompanying drawings.
[0048] The present invention relates to a method for discovering traffic abnormal data based on real-time big data analysis, and its flow chart is as Figure 1 shown, including the following steps:
[0049] I. Traffic multi-dimensional data collection step: Real-time collect a large amount of gantry transaction data, entrance and exit transaction data, and gantry license plate recognition data generated by expressways through Flume, and synchronize the real-time collected big data to different Topics corresponding to Kafka respectively according to the data type or collection source.
[0050] Mainly collect the entrance transaction flow of the entrance station, the exit transaction flow of the exit station, the gantry background gantry (valid / invalid) transaction flow, the gantry background gantry license plate recognition flow, the entrance station license plate recognition flow, the exit station license plate recognition flow, the entrance station virtual gantry flow, and the exit station virtual gantry flow.
[0051] For the transaction flow, it mainly includes: license plate (including license plate color), transaction amount, transaction time, transaction ID (required to be unique), passing PassId (required to be unique), passing medium, charging method, payment method, transaction address (toll station code, toll lane code, gantry code).
[0052] For the license plate recognition flow, it mainly includes: license plate (including license plate color), capture time.
[0053] Using Kafka as the message middleware can achieve more high-throughput and efficient asynchronous data processing, such as Figure 2The working principle of the Kafka message queue shown involves a Producer, a Kafka cluster (Kafka broker), a Consumer, and a ZooKeeper cluster (ZooKeeper Cluster). Among them, the Producer uses the push mode to send messages to the Kafka cluster. In practical applications, such as in the highway data collection scenario, this traffic multi-dimensional data collection step is equivalent to the Producer, which will actively push the collected messages such as gantry transaction data and entrance / exit transaction data to the Kafka cluster. The Kafka cluster consists of multiple Kafka brokers, which are the core nodes of the Kafka cluster and are responsible for receiving, storing, and forwarding messages. The messages pushed by the Producer will be stored in the Kafka broker. A Kafka broker can store messages of multiple Topics. In the Kafka cluster, multiple brokers work together to provide message storage and services for each Topic. In the highway data scenario, each collection point will send messages such as transaction data. The present invention will, according to the collection source, set different types of data such as "gantry transaction data" and "entrance / exit transaction data" as different Kafka Topics respectively. In this way, it is convenient to manage and process various messages. These messages classified by Topic will be scattered and stored on multiple Kafka brokers, thus realizing the efficient storage and management of data. The Consumer uses the pull mode to subscribe to and consume messages from the Kafka cluster. For example, in the subsequent data processing process, the steps responsible for data cleaning, analysis, etc. are the Consumer, and they will pull the stored highway transaction data from the Kafka cluster for subsequent processing operations. The ZooKeeper cluster (or ZK management cluster) is used to manage the Kafka cluster configuration: record the relevant configuration information of each broker in the Kafka cluster; elect a leader: in the Kafka cluster, when a certain broker fails or a new broker joins, ZooKeeper will be responsible for electing a new leader node. For example, when a certain Kafka broker recovers after a short failure, ZooKeeper will coordinate the election to ensure the normal operation of the cluster; load balancing: ZooKeeper can monitor the load conditions of each broker in the Kafka cluster and perform load balancing adjustments. If a certain broker has too high a load, ZooKeeper will assist in distributing some messages to other brokers with lower loads to ensure the overall performance stability of the cluster. Therefore Figure 2It is also an architecture that realizes efficient message transmission and processing through Kafka. The Producer sends messages, the Kafka cluster stores messages, the Consumer pulls messages for processing, and ZooKeeper manages to ensure the stable operation of the Kafka cluster.
[0054] II. Steps for intelligent hierarchical parsing of traffic data: Based on Flink, a data hierarchical architecture is constructed for intelligent hierarchical parsing of traffic data. This hierarchical architecture can make the data system clearer and simplify complex problems. As Figure 3 shown, this data hierarchical architecture can be understood as a data warehouse, which is divided into multiple levels. The data of each level is described as follows:
[0055] ① The original data layer (ODS, Operation Data Store) stores the unprocessed original data, which is consistent with the source system in structure and is the data preparation area of the data warehouse; that is, it stores the original big data collected in real time. The data is stored in Kafka and consists of the most original gantry transactions, entrance and exit water flows, and license plate recognition data.
[0056] ② The common dimension layer (DIM, Dimension) is constructed based on the dimension modeling theory and stores the dimension tables in the dimension model to save consistent dimension information. Specifically, several dimension tables including the gantry association information table, gantry basic information table, holiday information table, and vehicle type information table are constructed, and the corresponding dimension tables are pre-loaded into the Redis cache according to the scheduled task or trigger instruction.
[0057] Furthermore, the gantry association information table includes several field combinations such as gantry number, superior gantry number, inferior gantry number, section identification, and geographical coordinates; the gantry basic information table includes several field combinations such as gantry number, gantry name, toll type, billing rule version, and equipment status; the holiday information table includes several field combinations such as holiday date range, holiday type, free vehicle type category, and free policy description; the vehicle type information table includes several field combinations such as vehicle type code, vehicle type name, and toll coefficient.
[0058] The data is stored in the mysql relational database. Because the amount of this type of data is small and the changes are small, when the service starts, the full amount of data is directly imported into the redis cache, and when there are changes, the cache is directly updated.
[0059] ③ The detailed data layer (DWD, Data Warehouse Detail) is constructed based on the dimensional modeling theory, stores the fact tables in the dimensional model, and saves the operation records at the smallest granularity of each business process. Specifically, Flink is used to clean the raw big data collected in real time, and the cleaned data is associated with the corresponding dimension tables in the DIM layer through specified fields to generate standardized passing data.
[0060] After cleaning the raw data, it is associated with the dimension table. After dimension degradation, more refined and standardized minimum-granularity passing data is generated and stored in Kafka. It mainly includes the following fields: license plate (including color), generation time, reception time, whether it is a holiday, involved amount, involved gantry or entrance / exit station, involved owner, transaction ID, passing passid, passing medium, billing method, data type. For data that cannot be filled in for license plate recognition, fill in the blanks.
[0061] Among them, the data type can be further refined into: provincial boundary exit gantry data, provincial boundary license plate recognition data, virtual gantry exit gantry transaction data, exit station exit transaction data, exit station license plate recognition data, provincial boundary entrance gantry transaction data, provincial boundary entrance license plate recognition data, virtual gantry entrance gantry data, entrance station entrance transaction data, entrance station entrance license plate recognition data, ordinary gantry transaction data, ordinary gantry license plate recognition data.
[0062] Furthermore, cleaning the raw big data collected in real time can include uniformly converting the timestamps of different collection sources into Beijing time format through the Flink parallel computing engine, filtering records with negative or transaction amounts exceeding the preset threshold, completing or discarding records with missing gantry numbers, and judging and removing duplicate records based on the transaction serial number or timestamp. The cleaned data is associated with the corresponding dimension tables in the DIM layer through specified fields. Specifically, it can be: associating the gantry number field in the cleaned data with the gantry number fields in the DIM layer gantry association information table and the gantry basic information table respectively, associating the transaction time field in the cleaned data with the holiday date range field in the DIM layer holiday information table, and associating the vehicle type code field in the cleaned data with the vehicle type code field in the DIM layer vehicle type information table.
[0063] ④ The summary data layer (DWS, Data Warehouse Summary) is based on the upper-layer indicator requirements, takes the analysis theme object as the modeling driver, and constructs summary tables with common statistical granularity. This DWS layer processes data in three stages:
[0064] a. The first stage
[0065] Fusion stage: Read the standardized passing data generated in the DWD layer from Kafka, associate them by license plate, cut the multiple passes of the same vehicle into multiple segments of passing data according to the time sequence and passing logic, and then publish the processed multiple segments of passing data to Kafka.
[0066] Specifically, first read the smallest granular passing data in the DWD layer from Kafka, associate them according to the same license plate, then read the historical associated data in the redis cache or doris disk, perform fusion and deduplication, then sort by generation time, cut the entrance and exit data, and finally split the multiple passes of a vehicle into multiple segments and publish them to Kafka. The data structure after fusion contains the following main fields (license plate, passing passids, last transaction id, first transaction id, last transaction time, first transaction time, total amount received in the province at the exit, preferential amount in the province at the exit, cumulative total amount received at the gantry, type of the first passing data, type of the last passing data, owner associated with the first passing, owner associated with the last passing, set of passing data).
[0067] b. The second stage
[0068] Analysis stage: First read the data of the first stage from Kafka, then analyze item by item according to the exception indicators, and finally store the analysis results and the data after fusion into doris, cache the abnormal data into redis, and at the same time push the statistical relevant information to Kafka.
[0069] Specifically, read the multiple segments of passing data published in the fusion stage from Kafka, and based on the gantry association information table and the gantry basic information table in the DIM layer, perform real-time analysis of exception indicators including no exit for the entrance, inconsistent cumulative amount at the gantry and the amount at the exit, abnormal path, isolated point data, gantry time jump, license plate recognition without transaction, virtual gantry data but no entrance and exit transaction data, OBU or CPC cumulative amount billing but no gantry transaction data, license plate recognition error, license plate reverse label, license plate mislabel, etc. through the Flink real-time data stream processing engine, and publish the analysis results to Kafka and store them in the Doris database at the same time; during the real-time analysis process, based on the holiday information in the holiday information table in the DIM layer and the vehicle type data information in the vehicle type information table, automatically identify whether it is a holiday free vehicle type or free flow. If the abnormal data is generated by a holiday free vehicle type, it is automatically marked as a special exception, and if it is free flow, the flow is automatically determined to be normal.
[0070] c. The third stage
[0071] Statistics stage: First read the data of the second stage from Kafka, then aggregate and count according to the exception indicators, and finally store the statistical results in mysql and push them to the front end through the interface for real-time display.
[0072] Specifically, based on the analysis results pushed by the analysis stage received by the Flink node, aggregation statistics are performed according to preset dimensions to generate a vehicle anomaly details dataset, anomaly information, statistical information for each dimension, a details dataset after aggregation of a single vehicle passage, and analysis results. Further, the preset dimensions may include anomaly types, time windows, toll stations, etc., where the time window includes a time window in units of hours / days / weeks; the generated vehicle anomaly details dataset and anomaly information include the anomaly occurrence time, location, type, vehicle identification involved, etc., the statistical information for each dimension includes statistical results aggregated by time window, anomaly type, and toll station dimensions, etc., and the details dataset and analysis results after aggregation of a single vehicle passage include the complete passage path, billing details, etc.
[0073] ⑤ The data application layer (ADS, Application Data Service) stores the results of various statistical indicators and is a layer facing the end users and applications. Here, the data processed and analyzed by the previous layers is presented in an intuitive form to provide support for business decision-making and operation management.
[0074] The present invention uses Flink as the analysis framework to connect the subscription and consumption of data. At the same time, stream processing can process data in real time and in a rolling manner, and aggregation analysis can statistically process data more quickly and conveniently. As Figure 4 shown, it shows the working principle of Flink, which mainly involves Flink Program, Client, JobManager, and TaskManager, as follows:
[0075] Flink Program Code Writing: Developers write Flink program code to define data processing logic. For example, in the scenario of processing highway anomaly data, code for cleaning and analyzing gantry transaction data, entrance and exit transaction data, etc. will be written. The written code generates a Dataflowgraph through the Optimizer / Graph Builder, which is a graphical representation of the data processing logic, describing the data flow and each processing step. The Client submits the generated Dataflow graph to the JobManager, and this process is called Submitjob, which is equivalent to telling the JobManager that there is a new data processing task to execute. The Client also receives status updates from the JobManager, such as the execution progress of the job and whether there are any exceptions, etc., so that developers can understand the job execution situation. The JobManager is the master node of the Flink cluster, responsible for coordinating and managing the entire job execution. After receiving the job submitted by the Client, it schedules the job through the Scheduler, triggers the checkpoint operation through the Checkpoint Coordinator. The checkpoint mechanism is used to ensure that the job state can be restored in case of failures, guarantee data processing consistency, and receives Task Status, Heartbeats, and Statistics from the TaskManager to monitor the execution of each task. If a problem is found with a certain TaskManager, corresponding actions can be taken, such as reassigning tasks. The TaskManager is the worker node, which contains multiple Task Slots, and each Task Slot can run a Task. The TaskManager receives tasks from the JobManager and executes them, such as performing specific cleaning and calculation operations on highway data. It also manages memory and I / O resources through the Memory&I / OManager to ensure that resources are reasonably used during task execution, and is responsible for network communication through the Network Manager to transfer data streams with other TaskManagers. For example, when processing distributed data, different TaskManagers may need to exchange intermediate results, and feedback task status, heartbeats, and statistical information to the JobManager to let the JobManager know the task execution status in real time.In summary, through the collaborative work of the above components, Flink achieves efficient data processing in a distributed environment. From program writing and submission to job scheduling and execution, and then to resource management and status monitoring, it ensures the smooth progress of the entire data processing process.
[0076] The present invention applies Flink to the real-time analysis scenario of highway traffic abnormal data. Utilizing its distributed stream processing ability, it performs real-time processing on multi-source heterogeneous data such as gantry transactions, entrance and exit transactions, and gantry license plate recognition, and adapts according to the data characteristics and business requirements of the transportation industry (such as real-time discovery of abnormal data, accurate billing, etc.). Based on the Flink framework, a data hierarchical architecture (ODS, DWD, DIM, DWS, etc.) is constructed, and data cleaning, dimension association, abnormal analysis, etc. are performed using Flink at different levels. For example, in the DWD layer, Flink is used to clean the original traffic data, and in the DWS layer, abnormal index analysis is performed, etc., to organically combine the business-specific data processing logic with the Flink process.
[0077] III. Multi-dimensional output steps of traffic abnormal data: Provide the data generated in the statistical stage through the API interface, supporting multi-dimensional real-time query and warning.
[0078] Specifically, provide the vehicle abnormal detail data set and abnormal information, various dimension statistical information, as well as the detail data set and analysis results after vehicle single-pass aggregation generated in the statistical stage through the API interface. This API interface adopts load balancing technology and data caching strategy, supports communication protocols such as HTTP / HTTPS / RPC, and supports two data interaction modes of active push and passive pull, and supports multi-dimensional real-time query and warning, and is compatible with the Web side, mobile side, and third-party systems. Users can view the statistical results and abnormal details of this method in real time through the web side / mobile side / third-party system.
[0079] Compared with the prior art, the method of the present invention can more timely discover and handle problems. Starting from the upload of the first node data of vehicle passage, it breaks the tradition of waiting for N days and then analyzing in the reconciliation and auditing links. In addition, for data association, through the technical means of big data, both the efficiency and performance have been greatly improved, and it also provides more reliable and efficient data support for the subsequent auxiliary applications of the system. The positioning of this method is to discover and expose problems as quickly and as much as possible, and some data that are not problems may be classified as abnormal situations. For example, the definition of abnormal situations for in-transit vehicles and un-uploaded exits. If all in-transit vehicles are defined as un-uploaded exits, then the anomalies processed by the owners and maintenance personnel will cover all data. If in-transit vehicles are defined as normal, then how to define the situation of un-uploaded exits. This solution can set multiple time thresholds for anomaly warning in the second stage - the analysis stage of the DWS layer, and set the anomaly levels as possibly abnormal, slightly abnormal, moderately abnormal, and severely abnormal. As time goes by, the anomalies initially warned may become normal with the integrity of data upload. The DWS layer aims to provide efficient and fast summary data access for business analysis. By constructing a summary metric fact table with common granularity, it helps to identify trends, patterns, and anomalies. The analysis stage itself involves in-depth analysis of data to discover anomalies. Adding the time threshold anomaly warning logic can better judge data anomaly situations based on the existing data processing and analysis processes, combined with the time dimension, which conforms to the functional positioning of this stage. The present invention greatly reduces the data range for the audit and reconciliation system to analyze data, greatly relieves the pressure on the subsequent system to integrate and analyze data, and also reduces the pressure on the ministry-level system to access data, greatly improving the performance of the subsequent system.
[0080] The present invention realizes real-time analysis of traffic data such as gantries, license plate recognition, and entrance and exit flows generated on highways through big data methods, and can classify and pre-warn vehicle passage data suspected of anomalies. Specifically,
[0081] Classification pre-warning: In the analysis stage of the DWS layer, based on a variety of preset anomaly indicators for analysis, different anomaly indicators correspond to different types of abnormal data. For example, "having an entrance but no exit" is classified as an incomplete passage record type; "the cumulative amount of the gantry is inconsistent with the exit amount" is classified as an abnormal fee settlement type. By identifying these different types of anomalies, the abnormal data is classified and marked, and targeted prompts can be sent to relevant personnel or systems according to different types during pre-warning, facilitating quick positioning of the problem category.
[0082] Level-based warning: The level is determined based on factors such as the severity of abnormal data and the possible losses. For example, for the abnormal situation of "entry without exit" that may cause significant toll losses and has not been processed for a long time, it is set as a high-level warning; for some minor license plate recognition errors that may have little impact and occur occasionally, it is set as a low-level warning. During the warning process, the level difference is reflected through different notification methods (such as text messages, system pop-ups, etc.) or notification objects (managers at different levels), so that relevant personnel can give priority to handling high-level abnormalities and allocate resources reasonably.
[0083] The method for discovering traffic abnormal data based on real-time big data analysis in the present invention conducts real-time pre-analysis of traffic data through big data means, continuously integrates vehicle passing data during the analysis, and finally provides the integrated data and abnormal data to subsequent system applications. The specific implementation method is data collection → data analysis → data application, that is, corresponding to the above three steps (traffic multi-dimensional data collection step, traffic data intelligent hierarchical parsing step, and traffic abnormal data multi-dimensional output step). As Figure 5 shown in the preferred process.
[0084] 1. Data collection: mainly collect the entrance and exit data of toll stations, entrance and exit license plate recognition data, gantry transaction data, gantry license plate recognition data, such as Figure 5 shown in the data generation source. Each entrance and exit toll station and gantry equipment on the highway generate various transaction flow data and license plate recognition flow data, including entrance virtual gantry flow, entrance transaction flow, entrance license plate recognition flow, exit virtual gantry flow, exit transaction flow, exit license plate recognition flow, gantry transaction flow, gantry license plate recognition flow, etc. These data cover key information such as license plate, transaction time, amount, and passing medium. The timeliness of the upload of these data affects the timeliness of the subsequent entire analysis. Therefore, the original timed transmission mode in this aspect has been changed to real-time transmission. At the same time, the rust language is used for development, greatly improving the transmission efficiency and security. After the data is collected, it is directly synchronized to kafka by flume according to the theme.
[0085] Figure 5The data transmission pre - processing links are also shown, such as gateways and firewalls: All kinds of streaming data first flow to the gateway and the firewall. The gateway plays roles such as protocol conversion and data adaptation, ensuring that data generated by different devices can be transmitted in a suitable form. The firewall filters data for security, intercepting illegal or malicious data to prevent external attacks and data leakage, ensuring the security of data transmission. nginx: The data that has been preliminarily processed by the gateway and the firewall reaches nginx. nginx mainly undertakes the functions of reverse proxy and load balancing. It can reasonably distribute requests to different backend servers, avoid overloading a single server, improve the overall performance and stability of the system, and can also perform preliminary caching and forwarding of data.
[0086] The data reception and intermediate storage links are also shown, such as stream reception: The data processed by nginx is received by stream reception and integrates various streaming data forwarded from nginx, preparing for subsequent data processing. Kafka: The data integrated by stream reception is sent to Kafka. As a message middleware, Kafka can cache and transmit data, enabling asynchronous processing of data. It can decouple the data production and consumption links, ensuring stable operation even when the data generation rate and processing rate are inconsistent. At the same time, the high - throughput feature of Kafka is also suitable for processing a large amount of real - time data generated by highways.
[0087] The core is big data analysis and processing based on Flink, that is, data analysis is carried out by the constructed data hierarchical architecture (ODS layer, DWD layer, DIM layer, DWS layer, ADS layer). At the same time, database storage supports Redis, MySQL, and Doris. After big data analysis and processing by Flink, the final result data is transmitted to the front - end for display, and the processed data can be presented to users in visual forms such as reports and charts, facilitating managers to view and analyze, and assisting them in making decisions.
[0088] 2. Data Analysis: From the perspective of data flow, it refers to the process of taking the collected raw data through (1) data cleaning, (2) dimension association, (3) dimension degradation, (4) fusion aggregation, (5) status deduplication, (6) aggregation analysis, (7) grouped statistics, and (8) storage and landing. Among them, (1) Data Cleaning: Verify the accuracy of fields. For abnormal data such as license plate garbled characters, license plate colors not meeting specifications, abnormal passids in passages, and negative amounts, directly clean and output them. After the transaction is repaired, re - analyze; (2) Dimension Association: Associate the transaction with the dimension table, mainly by associating the basic information of the gantry, which is more convenient for us to fit the path, find the suspicious points in the path, and reverse - label and mis - label; (3) Dimension Degradation: After associating the data with the dimension table, directly bind the owner ID to the data, which is convenient for statistical analysis of various indicators according to the owner and reduces table association during statistics; (4) Fusion Aggregation: First, fuse the data with the same license plate based on the license plate. After status deduplication, finally, cut the data into multiple segments based on the entrance and exit times (in the case of a vehicle passing through multiple times); (5) Status Deduplication: During the transmission process, due to network fluctuations or human factors, the same data may be uploaded multiple times. Before data aggregation, it is necessary to cover and deduplicate according to the data primary key information, and use the latest and non - repeated data for fusion analysis. At the same time, for the definition of entrances and exits, there are multiple types of data. For entrances, there are entrance stop identifications, entrance provincial gantry identifications, entrance stop transactions, entrance stop virtual gantry transactions, and entrance provincial gantry transactions. For exits, there are exit stop identifications, exit stop transactions, exit stop virtual gantry transactions, exit provincial plate identifications, and exit provincial gantry transactions. Here, it is also necessary to deduplicate the entrance and exit transaction information and the identification information with similar times according to the principle that transactions are greater than identifications. When both entrance and exit transactions and virtual gantry transactions exist, do not deduplicate, and use the entrance and exit transactions as the start and end points of the path; (6) Aggregation Analysis and (7) Grouped Statistics, that is, the three stages of DWS (fusion stage, analysis stage, and statistics stage), as Figure 6 shown in the Flink process and Figure 7 the specific implementation example process shown. When missing data is found, the missing part will be retrieved from the historical analysis of abnormal data based on the passage passid, and after re - fusion aggregation and status deduplication, re - fusion analysis will be carried out.
[0089] Such as Figure 6As shown, after the process starts, the Flink node pulls Kafka data. As a distributed stream processing framework, Flink pulls data from the Kafka message middleware. The data stored in Kafka has a wide range of sources and can be regarded as a staging area for ODS layer (raw data layer) data, including unprocessed raw transaction flows and license plate recognition flows collected from various highway stations. The Flink node performs format verification on the pulled data. If the data format is normal, it enters the subsequent processing flow. If the data format is abnormal, the abnormal data is output to the Kafka abnormal stream. For data with a normal format, it is formatted and streamlined, converted into standardized minimum-granularity passing data, that is, the operation of the DWD layer, where the raw data is cleaned and converted into more standardized and finer-granularity processable data, and then output to Kafka. Then comes the DWS layer fusion stage. The data that has been formatted and streamlined is pulled from Kafka again, and it is checked whether the current vehicle association dataset exists in the Redis cluster. Redis is often used to cache dimension table data and some frequently accessed historical data, which plays a role in accelerating data query here. If it exists, the data in Redis and the current data are fused to construct pass_record. If it does not exist, it enters the Doris query step. If the relevant data is not found in Redis, the aggregated data (pass_record) for the vehicle within N days before and after the current time is queried from the Doris database table. As an MPP architecture database, Doris is suitable for storing and querying a large amount of analysis result data. If the relevant data exists, proceed to the next step. If it does not exist, construct the fused data (pass_record). If the relevant data is found in Redis or Doris, the current data and historical data are fused to construct new fused data (pass_record). For example, the current passing data of the vehicle and historical passing data are integrated according to certain rules. Then comes the DWS layer analysis stage, where the constructed fused data is analyzed item by item according to preset abnormal indicators (such as no exit, inconsistent amount, etc.). This is the core operation of the DWS layer analysis stage, and the analysis process is as Figure 7 shown. The fused data is flushed into Redis for caching to facilitate subsequent rapid access. The analysis results are stored in the Doris wide table and flushed into Kafka. Here, Kafka plays a role in message passing, passing the analysis result data to the subsequent processing flow. Then comes the DWS layer statistics stage. The Flink node receives the data, aggregates and statistics according to different indicator standards, outputs the statistical results and details, and performs further aggregation calculations on the analyzed data. Finally, the statistical results and details are pushed to the interface end for display to the user, completing the entire data processing and presentation process.
[0090] Such as Figure 7As shown in the figure, during the DWS layer analysis phase, by comprehensively analyzing the data related to highways, various abnormal situations are identified, and special processing is carried out for special situations (such as holidays, free flows, etc.) to improve data accuracy and analysis efficiency and reduce labor costs. The following types of abnormal situations are mainly analyzed: having an entrance but no exit; having a license plate identification but no transaction; isolated point data; the cumulative amount of gantries not matching the exit amount; abnormal path; having virtual gantry data but no entrance and exit transaction data; the OBU or CPC cumulative amount is charged, but there is no gantry transaction data; gantry time jump, license plate identification error, reverse label, mislabel. For example Figure 7 Judge in turn whether it is an isolated point, whether there is an exit, whether the exit amount is 0, whether there is an entrance, whether the exit is charged by the OBU or CPC cumulative amount, verify whether the amount of the normal transaction gantry data matches the exit amount, whether the cumulative amount of the gantry is greater than the exit amount, and whether there is isolated point data in redis according to the license plate and passid, etc.
[0091] Having an entrance but no exit: Judge whether there is a situation of having an entrance but no exit. If it exists, relevant information may be recorded first and then processed according to specific rules. When missing exit data is found, the missing part is retrieved from the historical analysis of abnormal data based on the passing passid. After re-fusion aggregation and status de-duplication, fusion analysis is carried out again. If the missing part is not found, the abnormal data may be marked or processed according to the established rules.
[0092] Having a license plate identification but no transaction: Judge whether there is a situation of having a license plate identification but no transaction. If it exists, relevant records will also be made and then further processed according to the situation. For example, check whether transaction information can be supplemented from other associated data. If not, the abnormal data will be processed according to the rules.
[0093] Isolated point data: Judge and process isolated point data. Isolated point data is a data point with extremely low correlation with other data. If isolated point data is detected, it will be analyzed whether it is valid data. If it is invalid, it may be deleted or marked; if it may be valid but lacks association, try to find associated information from historical data, and the processing method is similar to the above abnormal situations.
[0094] The cumulative amount of gantries not matching the exit amount: Check whether the cumulative amount of gantries is consistent with the exit amount. If not, analyze the reason for the difference, which may be done by looking up relevant transaction records and checking the charging rules, etc. If it is found that the difference is caused by missing data, the missing part is retrieved from the historical analysis of abnormal data based on the passing passid, and re-fusion aggregation, status de-duplication and fusion analysis are carried out again; if the difference is caused by special situations of charging rules, etc., it will be processed and recorded according to the corresponding rules.
[0095] Path anomaly: Determine whether the vehicle passing path is abnormal. If the path is abnormal, check whether it can be corrected by supplementing path information (such as obtaining from historical data or other relevant data sources). If it can be supplemented, re-perform operations such as fusion and aggregation; if it cannot be supplemented or corrected, mark the abnormal data.
[0096] There is virtual gantry data but no access transaction data: When this situation is detected, try to find the reason for the missing access transaction data. If the missing access transaction data can be retrieved from historical data, after retrieving it based on the passing passid, re-perform fusion and aggregation, status de-duplication, and fusion analysis; if it cannot be retrieved, process the abnormal data according to the rules.
[0097] The OBU or CPC accumulative amount is charged, but there is no gantry transaction data: When this anomaly occurs, analyze the reason for the lack of gantry transaction data. If the gantry transaction data can be supplemented from other data sources, after supplementation, re-perform operations such as fusion and aggregation; if it cannot be supplemented, mark or perform other processing on this abnormal data.
[0098] Gantry time jump, license plate recognition error, reverse label, mislabel: Judge and process situations such as gantry time jump, license plate recognition error, reverse label, mislabel. For example, when there is a license plate recognition error, try to assist in judging the correct license plate through other information (such as vehicle type, entrance information, etc.). If it can be corrected, re-process the relevant data; if it cannot be corrected, mark the abnormal data.
[0099] Holiday exception: Because type-I passenger cars are free on holidays, many toll stations directly let vehicles pass through at the entrances and exits without generating any transaction records, but the gantries are normally charged, which will result in a large amount of data with no exit but gantry data. And this part of the data not only does not generate revenue, but also increases the complexity of system analysis. If this part of the data is pushed to the owner for item-by-item verification, it will bring a large amount of labor costs. For the abnormal data generated by free vehicle types on holidays (such as type-I passenger cars being free on holidays), the present invention no longer pushes it to the owner for verification, simplifies the analysis process, and reduces the labor cost.
[0100] Free transaction exception: For free transactions generated at the exit due to various reasons such as green channel, holidays, vaccine vehicles, etc., the present invention no longer analyzes anomalies in combination with gantry data, directly determines that the transaction is normal, simplifies the analysis process, and also reduces the subsequent labor verification cost.
[0101] This (6) aggregation analysis process can effectively identify anomalies in highway data, improve data quality, and at the same time reduce the system analysis complexity and labor verification cost by simplifying the analysis process for special situations through careful judgment and processing of various abnormal situations and reasonable exception settings for special situations.
[0102] (7) Group statistics. In this stage, data is grouped and statistically analyzed based on the owners, source points (toll stations, gantries), and anomalies. For the daily gantry isolated point data, after grouping based on the gantries, sorting is performed, and the gantries that generate more isolated point data daily can be obtained. By re-analyzing the data generated by these gantries, gantries with time jumps, unclear license plate recognition, and reverse labels can be analyzed. It provides data support for calibrating gantry equipment across the province and improving the accuracy of gantry data.
[0103] (8) Storage and landing: According to data queries and the hot and cold levels of applications, the final statistical results and anomaly details are output to MySQL. For analysis details and results, they are all stored in Doris. Doris also adopts a hot and cold data storage model, transferring data after N days to cold data for storage, always ensuring the query efficiency of hot data.
[0104] 3. Data application: This method provides a web-based application to display anomaly data and details of various statistical dimensions, which are assigned to various personnel by role and permission, providing data query services for various personnel. This method can provide data support for systems such as provincial center auditing, reconciliation, online billing, and gantry summary upload. Subsequent systems no longer need to fuse all data.
[0105] The present invention also relates to a system for discovering traffic anomaly data based on real-time big data analysis. This system corresponds to the above-mentioned method for discovering traffic anomaly data based on real-time big data analysis and can be understood as a system for implementing the method for discovering traffic anomaly data based on real-time big data analysis. It includes a traffic multi-dimensional data acquisition module, a traffic data intelligent hierarchical parsing module, and a traffic anomaly data multi-dimensional output module connected in sequence. Each module works in coordination. Among them,
[0106] The traffic multi-dimensional data acquisition module: Real-time collects a large amount of gantry transaction data, entrance and exit transaction data, and gantry license plate recognition data generated on highways through Flume, and synchronizes the real-time collected big data to different Topics in Kafka respectively according to the data type or collection source;
[0107] The traffic data intelligent hierarchical parsing module: Builds a data hierarchical architecture based on Flink for intelligent hierarchical parsing of traffic data, including an ODS layer, a DIM layer, a DWD layer, and a DWS layer in sequence. Among them, the ODS layer: Stores the original big data collected in real-time; the DIM layer: Builds several dimension tables including a gantry association information table, a gantry basic information table, a holiday information table, and a vehicle type information table, and pre-loads the corresponding dimension tables into the Redis cache according to a scheduled task or a trigger instruction; the DWD layer: Cleans the original big data collected in real-time, and associates the cleaned data with the corresponding dimension tables in the DIM layer through specified fields to generate standardized traffic data; the DWS layer processes data in three stages:
[0108] a. Fusion stage: Read the standardized passing data generated at the DWD layer from Kafka, associate them according to the license plate, cut the multiple passes of the same vehicle into multiple segments of passing data in chronological order and passing logic, and then publish the processed multiple segments of passing data to Kafka;
[0109] b. Analysis stage: Read the multiple segments of passing data published in the fusion stage from Kafka. Based on the gantry association information table and gantry basic information table in the DIM layer, perform real-time analysis of abnormal indicators including entry without exit, inconsistent gantry cumulative amount and exit amount, path anomaly, and license plate recognition error through the Flink real-time data stream processing engine, and publish the analysis results to Kafka and store them in the Doris database at the same time; During the real-time analysis process, based on the holiday information in the holiday information table in the DIM layer and the vehicle type data information in the vehicle type information table, automatically identify whether it is a holiday free vehicle type or free flow. If the abnormal data is generated by a holiday free vehicle type, it is automatically marked as a special anomaly. If it is free flow, the flow is automatically determined to be normal;
[0110] c. Statistics stage: Based on the Flink node, receive the analysis results pushed in the analysis stage from Kafka, perform aggregation statistics according to the preset dimensions, and generate a vehicle anomaly detail data set and anomaly information, statistical information for each dimension, a detail data set and analysis results after aggregation of a single vehicle pass;
[0111] The traffic anomaly data multi-dimensional output module: Provide the data generated in the statistics stage externally through the API interface, supporting multi-dimensional real-time query and warning.
[0112] Figure 8This is the overall preferred architecture diagram of the system for discovering traffic anomaly data based on real-time big data analysis of the present invention. Its infrastructure layer includes a network, independent servers, operating systems, and middleware, which provide a basic operating environment for the entire system. The network ensures data transmission; the independent servers provide computing and storage resources; the operating system is the basic software for server operation; the middleware is used to support the collaborative operation between different software. The database layer uses MySQL, Doris, Redis, etc. as persistent storage, responsible for storing various types of data in the system. MySQL is suitable for traditional relational data storage; Doris (a database with an MPP architecture) is used to store data such as analysis results, supporting efficient big data query and analysis; Redis is used to cache frequently accessed data, such as dimension table data, temporary calculation results, etc., to improve data reading speed. The data layer includes stored procedures, data caching, custom functions, transactions, read and write databases, etc., which manage and operate on the data in the database. The traffic multi-dimensional data collection module is set in this data layer. The business layer mainly has a big data analysis platform (i.e., the traffic data intelligent hierarchical parsing module) with Yarn, Flink, Kafka, and Redis as the core, as well as a log recording and permission control module. Yarn: Responsible for cluster resource management and scheduling, allocating resources for computing tasks such as Flink. Flink: A distributed stream processing framework, that is, based on Flink to build a data hierarchical architecture for intelligent hierarchical parsing of traffic data, including the ODS layer, DIM layer, DWD layer, and DWS layer. Kafka: A message middleware, caching and transmitting data, realizing asynchronous processing of data, decoupling the data production and consumption links. Redis: In addition to caching data, it can also be used to store some temporary status data, etc. Log recording: Records the operations and events during the system operation, facilitating problem troubleshooting and auditing. Permission control: Manages users' access rights to system resources and functions, ensuring system security. The application layer includes WEB, which constructs the front-end interface using technologies such as Element Plus, Html5, CSS3, Vue3, and Ajax. The traffic anomaly data multi-dimensional output module is set in this application layer, providing an interactive interface for users to display the system processing results, such as data reports, visualization charts, etc., facilitating users to operate and view information, and also including business function services such as reconciliation, auditing, online billing, and gantry analysis.
[0113] Preferably, in the intelligent hierarchical parsing module of the traffic data, the gantry association information table constructed in the DIM layer includes several field combinations of gantry number, superior gantry number, inferior gantry number, affiliated road section identifier, and geographical coordinates. The gantry basic information table includes several field combinations of gantry number, gantry name, toll type, billing rule version, and equipment status. The holiday information table includes several field combinations of holiday date range, holiday type, free vehicle type category, and free policy description. The vehicle type information table includes several field combinations of vehicle type code, vehicle type name, and toll coefficient;
[0114] The DWD layer cleans the originally collected big data in real time, including uniformly converting the timestamps of different collection sources into Beijing time format through the Flink parallel computing engine, filtering records with negative or transaction amounts exceeding the preset threshold, completing or discarding records with missing gantry numbers, and judging and removing duplicate records based on the transaction serial number or timestamp; associating the cleaned data with the corresponding dimension tables in the DIM layer through specified fields. Specifically, the gantry number field in the cleaned data is associated with the gantry number fields in the gantry association information table and the gantry basic information table in the DIM layer respectively, the transaction time field in the cleaned data is associated with the holiday date range field in the holiday information table in the DIM layer, and the vehicle type code field in the cleaned data is associated with the vehicle type code field in the vehicle type information table in the DIM layer.
[0115] Preferably, in the intelligent hierarchical parsing module of the traffic data, during the analysis stage of the DWS layer, real-time analysis of abnormal indicators such as outlier data, gantry time jumps, license plate recognition without transactions, virtual gantry data without entrance / exit transaction data, OBU or CPC cumulative amount billing without gantry transaction data, license plate reverse marking, and license plate mislabeling is also performed through the Flink real-time data stream processing engine.
[0116] Preferably, in the intelligent hierarchical parsing module of the traffic data, during the statistical stage of the DWS layer, aggregation statistics are performed according to preset dimensions. The preset dimensions include abnormal type, time window, and toll station, where the time window includes time windows in units of hours / days / weeks; the generated vehicle abnormal detail data set and abnormal information include the time, location, type, and vehicle identification involved in the abnormality. The statistical information for each dimension includes the statistical results aggregated according to the time window, abnormal type, and toll station dimensions. The detail data set and analysis results aggregated for a single vehicle passage include the complete passage path and billing details.
[0117] It should be noted that the specific embodiments described above can enable those skilled in the art to understand the present invention more comprehensively, but do not limit the present invention in any way. Therefore, although this specification has described the present invention in detail with reference to the drawings and embodiments, those skilled in the art should understand that the present invention can still be modified or equivalently replaced. In short, all technical solutions and their improvements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the patent of the present invention.
Claims
1. A method for discovering abnormal traffic data based on real-time big data analysis, characterized in that, Including the following steps: Traffic multi-dimensional data collection step: Real-time collect a large amount of gantry transaction data, entrance and exit transaction data, and gantry license recognition data generated by expressways through Flume, and synchronize the real-time collected big data to different corresponding Topics in Kafka according to data types or collection sources; Traffic data intelligent hierarchical parsing step: Build a data hierarchical architecture based on Flink for intelligent hierarchical parsing of traffic data, including an ODS layer, a DIM layer, a DWD layer, and a DWS layer in sequence. Among them, the ODS layer: Store the original big data collected in real time; the DIM layer: Build several dimension tables including a gantry association information table, a gantry basic information table, a holiday information table, and a vehicle type information table, and pre-load the corresponding dimension tables into the Redis cache according to a scheduled task or a trigger instruction; the DWD layer: Clean the original big data collected in real time, and associate the cleaned data with the corresponding dimension tables in the DIM layer through specified fields to generate standardized passing data; the DWS layer processes the data in three stages: a. Fusion stage: Read the standardized passing data generated by the DWD layer from Kafka, associate them according to license plates, cut the multiple passes of the same vehicle into multiple pass data segments in chronological order and passing logic, and then publish the processed multiple pass data segments to Kafka; b. Analysis stage: Read the multiple pass data segments published in the fusion stage from Kafka, and based on the gantry association information table and the gantry basic information table in the DIM layer, perform real-time analysis of abnormal indicators including no exit at the entrance, inconsistent cumulative gantry amount and exit amount, path anomaly, and license recognition error through the Flink real-time data stream processing engine, and publish the analysis results to Kafka and store them in the Doris database at the same time; During the real-time analysis process, based on the holiday information in the holiday information table in the DIM layer and the vehicle type data information in the vehicle type information table, automatically identify whether it is a holiday free vehicle type or free flow. If the abnormal data is generated by a holiday free vehicle type, it is automatically marked as a special anomaly. If it is free flow, the flow is automatically determined to be normal; c. Statistics stage: Based on the Flink node, receive the analysis results pushed in the analysis stage from Kafka, perform aggregation statistics according to preset dimensions, and generate a vehicle abnormal detail data set and abnormal information, various dimension statistics information, a detail data set and analysis results after aggregation of a single vehicle pass; Traffic abnormal data multi-dimensional output step: Provide the data generated in the statistics stage externally through an API interface, supporting multi-dimensional real-time query and warning.
2. The method for discovering traffic anomaly data based on real-time big data analysis according to claim 1, characterized in that In the intelligent hierarchical parsing step of the traffic data, the gantry association information table constructed by the DIM layer includes several field combinations of gantry number, superior gantry number, inferior gantry number, section identification, and geographical coordinates. The gantry basic information table includes several field combinations of gantry number, gantry name, toll type, billing rule version, and equipment status. The holiday information table includes several field combinations of holiday date range, holiday type, free vehicle type category, and free policy description. The vehicle type information table includes several field combinations of vehicle type code, vehicle type name, and toll coefficient.
3. The method for discovering traffic anomaly data based on real-time big data analysis according to claim 2, characterized in that, In the intelligent hierarchical parsing step of the traffic data, the DWD layer cleans the originally collected big data in real time, including uniformly converting the timestamps of different collection sources into Beijing time format through the Flink parallel computing engine, filtering records with negative or transaction amounts exceeding the preset threshold, completing or discarding records with missing gantry numbers, and judging and removing duplicate records based on the transaction serial number or timestamp; associating the cleaned data with the corresponding dimension tables of the DIM layer through specified fields. Specifically, the gantry number field in the cleaned data is associated with the gantry number fields in the gantry association information table and the gantry basic information table of the DIM layer respectively, the transaction time field in the cleaned data is associated with the holiday date range field in the holiday information table of the DIM layer, and the vehicle type code field in the cleaned data is associated with the vehicle type code field in the vehicle type information table of the DIM layer.
4. The method for discovering traffic anomaly data based on real-time big data analysis according to any one of claims 1 to 3, characterized in that In the analysis stage of the DWS layer, real-time analysis of abnormal indicators such as outlier data, gantry time jumps, license plate recognition without transactions, virtual gantry data without entrance / exit transaction data, OBU or CPC cumulative amount billing without gantry transaction data, license plate reverse marking, and license plate mislabeling is also performed through the Flink real-time data stream processing engine.
5. The method for discovering traffic abnormal data based on real-time big data analysis according to any one of claims 1 to 3, characterized in that In the statistical stage of the DWS layer, aggregation statistics are performed according to preset dimensions. The preset dimensions include abnormal type, time window, and toll station, where the time window includes time windows in units of hours / days / weeks; the generated vehicle abnormal detail data set and abnormal information include the time, location, type, and vehicle identification involved in the abnormality. The statistical information for each dimension includes the statistical results aggregated by time window, abnormal type, and toll station dimensions. The detail data set and analysis results aggregated for a single vehicle passage include the complete passage path and billing details.
6. The method for discovering traffic anomaly data based on real-time big data analysis according to any one of claims 1 to 3, characterized in that, In the multi-dimensional output step of the traffic abnormal data, the vehicle abnormal detail data set and abnormal information, the statistical information for each dimension, and the detail data set and analysis results aggregated for a single vehicle passage generated in the statistical stage are provided externally through the API interface. The API interface adopts load balancing technology and data caching strategy, supports HTTP / HTTPS / RPC communication protocols, and supports two data interaction modes of active push and passive pull, and supports multi-dimensional real-time query and warning.
7. A system for discovering traffic abnormal data based on real-time big data analysis, characterized in that, It includes a traffic multi-dimensional data collection module, a traffic data intelligent hierarchical parsing module, and a traffic abnormal data multi-dimensional output module connected in sequence. The traffic multi-dimensional data collection module: Real-time collects a large amount of gantry transaction data, entrance and exit transaction data, and gantry license recognition data generated by highways through Flume, and synchronizes the real-time collected big data to different corresponding Topics in Kafka according to the data type or collection source; The traffic data intelligent hierarchical parsing module: Based on Flink, constructs a data hierarchical architecture for intelligent hierarchical parsing of traffic data, successively including the ODS layer, DIM layer, DWD layer, and DWS layer. Among them, the ODS layer: Stores the original big data collected in real time; the DIM layer: Constructs several dimension tables including a gantry association information table, a gantry basic information table, a holiday information table, and a vehicle type information table, and pre-loads the corresponding dimension tables to the Redis cache according to a scheduled task or a trigger instruction; the DWD layer: Cleans the original big data collected in real time, and associates the cleaned data with the corresponding dimension tables in the DIM layer through specified fields to generate standardized traffic data; the DWS layer processes the data in three stages: a. Fusion stage: Reads the standardized traffic data generated by the DWD layer from Kafka, associates it according to the license plate, cuts the multiple passages of the same vehicle into multiple passage data according to the time sequence and traffic logic, and then publishes the processed multiple passage data to Kafka; b. Analysis stage: Reads the multiple passage data published in the fusion stage from Kafka. Based on the gantry association information table and the gantry basic information table in the DIM layer, through the Flink real-time data stream processing engine, conducts real-time analysis of abnormal indicators including no exit at the entrance, inconsistent cumulative amount of the gantry and the exit amount, abnormal path, and license recognition error, and publishes the analysis results to Kafka and stores them in the Doris database at the same time; During the real-time analysis process, based on the holiday information in the holiday information table in the DIM layer and the vehicle type data information in the vehicle type information table, automatically identify whether it is a holiday free vehicle type or free flow. If the abnormal data is generated by a holiday free vehicle type, it is automatically marked as a special abnormality. If it is free flow, the flow is automatically determined to be normal; c. Statistics stage: Based on the Flink node, receives the analysis results pushed in the analysis stage from Kafka, aggregates and statistics according to the preset dimensions, and generates a vehicle abnormal detail data set and abnormal information, various dimension statistical information, a detail data set and analysis results after aggregation of a single vehicle passage; The traffic abnormal data multi-dimensional output module: Provides the data generated in the statistics stage externally through an API interface, supporting multi-dimensional real-time query and warning.
8. The system for discovering traffic abnormal data based on real-time big data analysis according to claim 7, wherein, In the intelligent hierarchical parsing module of traffic data, the gantry association information table constructed in the DIM layer includes several field combinations of gantry number, superior gantry number, inferior gantry number, section identification, and geographical coordinates. The gantry basic information table includes several field combinations of gantry number, gantry name, toll type, billing rule version, and equipment status. The holiday information table includes several field combinations of holiday date range, holiday type, free vehicle type category, and free policy description. The vehicle type information table includes several field combinations of vehicle type code, vehicle type name, and toll coefficient. In the DWD layer, the original big data collected in real time is cleaned, including uniformly converting the timestamps of different collection sources into Beijing time format through the Flink parallel computing engine, filtering records with negative or transaction amounts exceeding the preset threshold, completing or discarding records with missing gantry numbers, and judging and removing duplicate records based on the transaction serial number or timestamp. The cleaned data is associated with the corresponding dimension tables in the DIM layer through specified fields. Specifically, the gantry number field in the cleaned data is associated with the gantry number fields in the gantry association information table and the gantry basic information table in the DIM layer, the transaction time field in the cleaned data is associated with the holiday date range field in the holiday information table in the DIM layer, and the vehicle type code field in the cleaned data is associated with the vehicle type code field in the vehicle type information table in the DIM layer.
9. The system for discovering traffic anomaly data based on real-time big data analysis according to claim 7 or 8, characterized in that, In the intelligent hierarchical parsing module of traffic data, during the analysis stage of the DWS layer, real-time analysis of abnormal indicators such as outlier data, gantry time jumps, license plate recognition without transactions, virtual gantry data without entrance / exit transaction data, OBU or CPC cumulative amount billing without gantry transaction data, license plate reverse marking, and license plate mislabeling is also performed through the Flink real-time data stream processing engine.
10. The system for discovering traffic anomaly data based on real-time big data analysis according to claim 7 or 8, characterized in that, In the intelligent hierarchical parsing module of traffic data, during the statistical stage of the DWS layer, aggregation statistics are performed according to preset dimensions. The preset dimensions include abnormal type, time window, and toll station, where the time window includes time windows in units of hours / days / weeks. The generated vehicle abnormal detail data set and abnormal information include the time, location, type, and vehicle identification involved in the abnormality. The statistical information for each dimension includes the statistical results aggregated by time window, abnormal type, and toll station dimensions. The detail data set and analysis results after aggregation of a single vehicle passage include the complete passage path and billing details.
Citation Information
Cited By
Data auditing method and system based on real-time data warehouse hierarchical architecture, equipment and medium
CN121301489A
Campus intelligent consumption analysis system and method based on multi-dimensional data fusion
CN121304226A
Traffic free flow charging method and system based on state snapshot
CN122493545A
A traffic free-flow tolling method and system based on state snapshots
CN122493545B