High-speed auditing system based on data ETL processing

Through the high-speed audit system with data ETL processing, the problem of low intelligence in traditional audit systems is solved, efficient and accurate auditing is achieved, and the safety and economic operation of highways are ensured.

CN120277136APending Publication Date: 2025-07-08HEILONGJIANG LONGTONG INTELLIGENT ELECTRONIC TOLL OPERATION DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304451.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional audit systems have low intelligence and poor efficiency and accuracy, making it difficult to meet the needs of modern highway management.

Method used

A high-speed audit system based on data ETL processing is adopted to build a path and cost model through data source acquisition, analysis, retrieval and audit modules, and combine vehicle information comparison to automatically identify abnormal behaviors and generate warning information.

Benefits of technology

It improves the efficiency and accuracy of the audit system, can promptly detect abnormal behaviors, ensure the safety and economic benefits of highway operations, optimize resource utilization, and reduce manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277136A_ABST
    Figure CN120277136A_ABST
Patent Text Reader

Abstract

The invention discloses a high-speed auditing system based on data ETL processing, and the system comprises a data source collection module which is used for collecting the related information of a data source; the data source analysis module is used for analyzing the data source related information to obtain data source analysis information, and the data source analysis information comprises evaluation passing and evaluation non-passing; the data calling module is used for selecting a calling data source from the data sources passing the evaluation to carry out data calling and carrying out ETL (Extract Transform and Load) processing on the called data; the data auditing module is used for auditing the data subjected to ETL processing to generate warning information; and the information sending module is used for sending the generated warning information to a preset receiving terminal. According to the invention, more efficient and intelligent high-speed inspection can be carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audit systems, and particularly to a high-speed audit system based on data ETL processing. Background Art

[0002] With the continuous improvement of the highway network and the increasing traffic flow, highway management faces many challenges, such as the accurate collection of vehicle tolls, the precise monitoring of vehicle driving routes, and the timely discovery of abnormal vehicle behaviors. The traditional manual audit method is inefficient and inaccurate, and it is difficult to meet the needs of modern highway management. There is an urgent need for an efficient and accurate audit system to ensure the safety and benefits of highway operations.

[0003] Existing audit systems have low intelligence and mostly require manual audits, which has a certain impact on the use of audit systems. Therefore, a high-speed audit system based on data ETL processing is proposed. Summary of the Invention

[0004] In view of the deficiencies in the prior art, the present invention provides a high-speed audit system based on data ETL processing, including:

[0005] A data source collection module for collecting information related to the data source;

[0006] A data source analysis module for analyzing the information related to the data source to obtain data source analysis information, where the data source analysis information includes passing the evaluation and failing the evaluation;

[0007] A data retrieval module for selecting and retrieving data sources from the data sources that have passed the evaluation for data retrieval, and performing ETL processing on the retrieved data;

[0008] A data audit module for auditing the data processed by ETL to generate warning information;

[0009] An information sending module for sending the generated warning information to a preset receiving terminal.

[0010] Furthermore, the specific content of the data source collection module for collecting information related to the data source is as follows: collecting the information on the data source being retrieved and the information on the data source being adopted;

[0011] The specific process of analyzing the information related to the data source is as follows: extracting the information related to the data source, and obtaining the information on the data source being retrieved and the information on the data source being adopted from the information related to the data source;

[0012] The information on the data source being retrieved includes the number of times each source data source is retrieved and the time point of each retrieval;

[0013] The information that the data source is adopted is the number of times the data source is finally adopted;

[0014] Process the number of times each data source is retrieved, the time points of each retrieval, and the number of times the data source is finally adopted to obtain a retrieval evaluation parameter. When the retrieval evaluation parameter is greater than the preset value, it indicates that the evaluation passes.

[0015] Furthermore, the process of obtaining the retrieval evaluation parameter is as follows:

[0016] Mark the number of times each data source is retrieved as C, that is, the total number of times the data source is retrieved during the system operation;

[0017] Mark the time points of each retrieval of the data source as T=(t1, t2, t3... tC);

[0018] Mark the number of times the data source is finally adopted as A;

[0019] For data sources with the number of retrievals C>1, calculate the time interval between adjacent retrievals Δti = ti+1 - ti, i = 1, 2,..., C - 1;

[0020] After that, calculate the mean value of all Δti to obtain the average retrieval interval U;

[0021] Extract the first retrieval time and the last retrieval time of the number of retrievals C of the data source, calculate the difference between the last retrieval time and the first retrieval time to obtain the total retrieval time Tg;

[0022] Calculate the ratio of the total retrieval Tg to C to obtain the number of retrievals per unit time N;

[0023] After that, calculate the ratio of the number of times the data source is finally adopted A to the number of times the data source is retrieved C to obtain the adoption ratio Ac;

[0024] Finally, through the formula (U + N + Ac) / 3 = P, the retrieval evaluation parameter P is obtained.

[0025] Furthermore, the process of the data retrieval module selecting a data source from the data sources that pass the evaluation is as follows: Extract the retrieval evaluation parameters P of all data sources, and set the data sources corresponding to the two largest retrieval evaluation parameters P as the selected data sources.

[0026] Furthermore, extract the data sources where the retrieved evaluation parameter P is less than the preset value, mark them as abnormal data sources, and conduct regular monitoring on the abnormal data sources to monitor the changes in the retrieved evaluation parameter P of the abnormal data sources in the future, that is, regularly collect the retrieved evaluation parameter P of the abnormal data sources, plot the retrieved evaluation parameter P of the abnormal data sources into a line chart according to the collection time sequence, and then analyze the trend of the line chart. When the trend of the line chart is in a flat state or a downward state, generate data source abnormal information.

[0027] Furthermore, the generation process of the warning information is as follows:

[0028] First, construct an auditing model, which includes a path model and a cost model;

[0029] After that, collect vehicle information and compare the vehicle information with the path model and the cost model;

[0030] When any comparison is abnormal, generate warning information.

[0031] Furthermore, the construction process of the path model is as follows: Extract all the entrance stations and exit stations from the data after ETL processing;

[0032] After that, use a preset algorithm to calculate the shortest path from all entrance stations to exit stations, that is, obtain the path model. The preset algorithm includes Dijkstra algorithm or A* algorithm;

[0033] The construction process of the cost model is as follows;

[0034] Extract the unit mileage charging standards and the shortest paths of different vehicle types, mark the unit mileage charging standards of different vehicle types as Fi, mark the shortest path as G, and i is the number of vehicle types;

[0035] Through the formula G*Fi = Gf, that is, obtain the cost model Gf.

[0036] Furthermore, extract vehicle information, which includes vehicle type, vehicle entrance station, vehicle exit station and exit payment;

[0037] Extract the path model corresponding to the vehicle entrance station and vehicle exit station from all the path models;

[0038] After that, collect the actual driving mileage information of the vehicle through the road gantry, mark the actual driving mileage information as Z1, and mark the path model corresponding to the vehicle entrance station and vehicle exit station as Z2;

[0039] Through the formula (Z1 - Z2)*α = Zz, that is, obtain the evaluation coefficient. When the evaluation coefficient is greater than the preset value, it indicates that the comparison is abnormal and generate warning information. α is a correction value, 1.01 ≤ α ≤ 1.1, and α is proportional to Z2;

[0040] Extract the vehicle type, retrieve the standard cost per unit mileage corresponding to the vehicle type from the preset database, and mark it as E1;

[0041] Obtain the standard cost Ez through the formula E1 * Z1 = Ez, mark the outbound payment as E2, calculate the difference between E2 and E1, and obtain the cost difference Ee. When Ee exceeds the preset range, it indicates an abnormal comparison and a warning message is generated.

[0042] Furthermore, the content audited by the data audit module also includes abnormal vehicle behaviors. When abnormal vehicle behaviors are detected, a warning message is generated.

[0043] Furthermore, the determination process of abnormal vehicle behaviors is as follows:

[0044] Extract the historical passing records of the vehicle from the data after ETL processing; the historical passing records include license plate number, vehicle type, entry toll station, exit toll station, passing time (accurate to seconds or even milliseconds), toll amount, ETC transaction status, gantry capture time and location for each passing;

[0045] Associate and integrate the data from different data sources, including the lane toll system, gantry system, and ETC system, according to the vehicle identifier, i.e., the license plate number, to form a passing archive for each vehicle;

[0046] Calculate the number of passes of the vehicle within a preset time period, set the normal passing frequency range obtained through statistical analysis of historical data. When the number of passes of the vehicle exceeds the upper limit of this range, it is determined as abnormal and a warning message is generated. For example, under normal circumstances, a small passenger car passes through a certain section 3 - 5 times a week on average, and if a vehicle reaches more than 15 times and there is no reasonable explanation for its operation business (such as missing filing information for logistics distribution vehicles, etc.), it is marked as abnormal passing frequency.

[0047] Extract the passing interval of the vehicle from the vehicle passing archive, use time series analysis technology to model the passing interval time of the vehicle, such as using the ARIMA model, predict the reasonable next passing time interval of the vehicle. If the deviation between the actual passing interval and the predicted value exceeds the preset value, for example, the predicted interval is 2 days, but it actually passes again within 1 hour, it indicates that the vehicle's behavior does not conform to the normal passing rule and is determined as abnormal, and a warning message is generated.

[0048] Based on the historical traffic flow data of the highway, peak periods are divided (such as the morning and evening rush hours on weekdays, the outbound and return peak hours on holidays, etc.). The number of times a vehicle passes during the peak period and the total number of passes are statistically counted from the vehicle passing records. If the ratio of the number of times a vehicle passes during the peak period to the total number of passes exceeds a preset value (for example, more than 80%, and this ratio for normal vehicles may be between 30% - 50%), and through the analysis of the vehicle passing records, it is known that the vehicle is not an emergency rescue or bus special line vehicle, it is determined as abnormal and a warning message is generated;

[0049] Analyze the vehicle passing records to identify combinations of toll stations or road sections where the vehicle frequently travels back and forth within a distance less than the preset value, that is, set a distance threshold (such as within 50 kilometers). If the vehicle travels back and forth within this short-distance range multiple times (5 times or more) within the preset time period (such as within 3 consecutive days), and the fluctuation range of the toll amount for each time is less than the preset value, this may conform to the characteristics of card swapping for toll evasion or taking advantage of short-distance billing loopholes, and it is determined as abnormal and a warning message is generated;

[0050] For different vehicle types, statistically calculate the average toll level per kilometer on each road section, establish a vehicle type - road section toll benchmark model, and analyze the vehicle passing records. When it is found that within the preset time period (such as one month), the driving mileage of the vehicle is greater than the preset value, and the ratio of the normal toll amount of the same vehicle type to the total toll amount of the current vehicle is less than the preset value, and the application of the free passage policy is excluded, it is determined as abnormal and a warning message is generated.

[0051] The present invention has the following advantages compared with the prior art: This high-speed auditing system based on data ETL processing collects detailed data source retrieval and adoption information through the data source collection module. The data source analysis module accurately judges the value of the data source according to a rigorous retrieval evaluation parameter calculation method. For example, by comprehensively considering factors such as the retrieval frequency, time interval, and adoption ratio of the data source, high-quality data sources that pass the evaluation are screened out to ensure that subsequent data processing is based on reliable and frequently used data, improving the quality of the basic data for auditing and avoiding wasting resources on low-quality or irrelevant data sources.

[0052] The data retrieval module accurately selects the data source according to the retrieval evaluation parameter P, preferentially retrieves the most valuable data for ETL processing, optimizes the data retrieval process, improves the efficiency of data acquisition and processing, enables the entire auditing system to quickly focus on key data, and timely discovers potential problems.

[0053] For the data source, abnormalities can be detected in a timely manner. Data sources with a retrieval evaluation parameter P less than the preset value are marked as abnormal data sources and monitored regularly. By analyzing the trend through drawing a line chart, potential problems that may exist in the data source can be discovered in advance, such as the data stability deteriorating or gradually losing its reference value, so as to timely adjust the data source strategy and ensure the reliability of data supply.

[0054] During the auditing process, the constructed auditing model covers multiple aspects such as routes and fees. By accurately comparing with the actual vehicle information, whether it is an abnormal driving route or an abnormal toll fee, warning information can be quickly generated, effectively preventing illegal acts such as evading tolls and safeguarding the economic interests of highway operators.

[0055] Deeply analyze the abnormal behaviors of vehicles, integrate data from multiple systems to form vehicle passing files. Monitor from multiple dimensions such as passing frequency, interval time, passing proportion during peak hours, short-distance round-trip situations, and the rationality of vehicle type tolls. It can accurately identify vehicles with potential violations or abnormal operations such as high-frequency passing without reasonable reasons, abnormal passing times, illegal peak travel, short-distance abnormal round-trips, and unreasonable tolls, maintain the normal passing order of the highway, ensure fairness, and at the same time help discover potential safety hazards or clues of illegal operations, making this system more worthy of popularization and use. Brief Description of the Drawings

[0056] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. In all the drawings, similar elements or parts are generally marked with similar reference numerals. In the drawings, the elements or parts do not necessarily need to be drawn according to the actual ratio.

[0057] Figure 1 It is the system block diagram of the present invention. Detailed Embodiments

[0058] The following will describe in detail the embodiments of the technical solutions of the present invention in conjunction with the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, so they are only examples and cannot be used to limit the protection scope of the present invention.

[0059] It should be noted that unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meaning understood by those skilled in the art to which the present invention belongs.

[0060] As Figure 1 shown, a high-speed auditing system based on data ETL processing includes:

[0061] A data source collection module for collecting information related to the data source;

[0062] A data source analysis module for analyzing the information related to the data source to obtain data source analysis information, and the data source analysis information includes passing the evaluation and failing the evaluation;

[0063] A data retrieval module for selecting and retrieving data sources from the data sources that have passed the evaluation, and performing ETL processing on the retrieved data;

[0064] A data auditing module, used to audit the data processed by ETL to generate warning information;

[0065] An information sending module, used to send the generated warning information to a preset receiving terminal.

[0066] The specific content of the data source collection module for collecting data source-related information is as follows: collecting the information of the data source being retrieved and the information of the data source being adopted;

[0067] The specific process of analyzing the data source-related information is as follows: extracting the data source-related information, and obtaining the information of the data source being retrieved and the information of the data source being adopted from the data source-related information;

[0068] The information of the data source being retrieved includes the number of times each source data source is retrieved and the time point of each retrieval;

[0069] The information of the data source being adopted is the number of times the data source is finally adopted;

[0070] Processing the number of times each data source is retrieved, the time point of each retrieval, and the number of times the data source is finally adopted, to obtain a retrieval evaluation parameter. When the retrieval evaluation parameter is greater than a preset value, it means that its evaluation passes;

[0071] By collecting the information of the data source being retrieved and the information of the data source being adopted, it is possible to quantitatively evaluate the usage of different data sources. A data source with a large number of retrievals and a large number of final adoptions often means that its data quality is high and its relevance to the business is strong. For example, in a high-speed auditing system, if a certain data source is often retrieved and used in actual auditing work multiple times, it indicates that it can provide valuable information, and such a data source is more trustworthy. By calculating the retrieval evaluation parameter and comparing it with the preset value, these high-quality data sources can be screened out to ensure that subsequent data processing and analysis are based on a reliable data foundation.

[0072] Reducing interference from low-quality data: For those data sources with few retrievals and few adoptions, their data may have quality problems or do not match the business requirements. Through this evaluation process, these low-quality data sources can be excluded to avoid their interference with subsequent auditing work and improve the accuracy and efficiency of the system.

[0073] In a high-speed auditing system, there may be multiple data sources to choose from. If the data sources are not evaluated, data may be blindly retrieved from each data source, resulting in waste of resources. By calculating the retrieval evaluation parameter, it is possible to identify which data sources are more valuable, so as to preferentially retrieve data from these data sources, concentrate limited resources on high-quality data sources, and improve the efficiency of data retrieval.

[0074] Reducing the retrieval and processing of low-quality data sources can reduce the computing resource consumption and storage costs of the system. For example, avoiding ETL processing of a large amount of useless data can reduce the processing burden of the system, improve the operating efficiency of the system, and at the same time reduce the costs of hardware devices and software systems.

[0075] The process of obtaining the retrieval evaluation parameters is as follows:

[0076] Mark the number of times each data source is retrieved as C, that is, the total number of times the data source is retrieved during the operation of the system;

[0077] Mark the time points of each retrieval of the data source as T=(t1, t2, t3... tC);

[0078] Mark the number of times the data source is finally adopted as A;

[0079] For data sources with the number of retrievals C>1, calculate the time interval Δti=ti + 1 - ti between two adjacent retrievals, where i = 1, 2,..., C - 1;

[0080] After that, calculate the mean value of all Δti to obtain the average retrieval interval U;

[0081] Extract the first retrieval time and the last retrieval time of the number of retrievals C of the data source, calculate the difference between the last retrieval time and the first retrieval time to obtain the total retrieval time Tg;

[0082] Calculate the ratio of the total retrieval Tg to C to obtain the number of retrievals per unit time N;

[0083] After that, calculate the ratio of the number of times A the data source is finally adopted to the number of retrievals C of the data source to obtain the adoption ratio Ac;

[0084] Finally, through the formula (U + N + Ac) / 3 = P, that is, obtain the retrieval evaluation parameter P.

[0085] The process of the data retrieval module selecting data sources from the data sources that pass the evaluation is as follows: Extract the retrieval evaluation parameters P of all data sources, and set the data sources corresponding to the two largest retrieval evaluation parameters P as the selected data sources;

[0086] Considering both the usage frequency and timeliness comprehensively, by introducing three indicators: the average retrieval interval U, the number of retrievals per unit time N, and the adoption ratio Ac, the value of the data source is evaluated comprehensively and multi-dimensionally. The average retrieval interval U reflects the time pattern of data source retrieval. A shorter interval means that the data source may need to be updated or used frequently, and its timeliness requirement is higher. The number of retrievals per unit time N reflects the usage frequency of the data source. High-frequency usage indicates that the data source plays an important role in the business. The adoption ratio Ac measures the quality and availability of the data provided by the data source. A higher adoption ratio means that the data provided by the data source better meets the business requirements.

[0087] Avoid the limitations of single-index evaluation. A single index cannot comprehensively reflect the true value of the data source. For example, only considering the number of retrievals may select data sources that are retrieved frequently but have low data quality. And only focusing on the adoption ratio may ignore some data sources that are less adopted but very important in specific scenarios. By combining these three indicators and calculating the retrieval evaluation parameter P, the comprehensive value of the data source can be evaluated more accurately, avoiding wrong choices due to the limitations of single indicators.

[0088] Precisely select high-quality data sources. When retrieving data, choosing the two data sources with the largest retrieval evaluation parameter P can ensure the acquisition of the most valuable data. These data sources perform well in terms of usage frequency, timeliness, and data quality, providing a reliable data foundation for subsequent ETL processing and auditing work. For example, in a high-speed auditing system, selecting such data sources can more accurately detect abnormal data and violations, improving the accuracy and efficiency of auditing.

[0089] Reduce unnecessary data retrievals. By screening data sources through evaluation parameters, it is possible to avoid retrieving data from a large number of irrelevant or low-quality data sources, reducing unnecessary system overhead and time costs. The system can directly focus on the most valuable data sources, improving the efficiency of data retrieval and reducing the complexity of data processing.

[0090] Optimize system resource utilization. In the case of limited system resources, preferentially retrieving data from data sources with high evaluation parameters can allocate system resources more reasonably. These data sources can provide more valuable data, and investing resources in retrieval and processing can obtain higher returns. For example, for a high-speed auditing system with limited storage and computing resources, selecting high-quality data sources can improve the utilization efficiency of resources and avoid resource waste.

[0091] Improve the stability of data supply: Selecting data sources with high evaluation parameters means that these data sources have performed stably in historical usage and can provide the required data on time and accurately. This helps to improve the stability of the system's data supply and reduce system failures and business interruptions caused by data missing or errors.

[0092] Enhance the anti-interference ability of the system: The data quality of high-quality data sources is relatively high, which can reduce the interference and errors caused by low-quality data. In a complex business environment, the system can better handle various data changes and anomalies, ensuring the stable operation of the system.

[0093] Extract the data sources whose retrieved evaluation parameter P is less than the preset value, mark them as abnormal data sources, and conduct regular monitoring on the abnormal data sources to monitor the changes in the retrieved evaluation parameter P of the abnormal data sources in the future, that is, regularly collect the retrieved evaluation parameter P of the abnormal data sources, plot the retrieved evaluation parameter P of the abnormal data sources into a line chart according to the collection time sequence, and then analyze the trend of the line chart. When the trend of the line chart is in a flat state or a downward state, generate data source abnormal information;

[0094] By extracting the data sources whose retrieved evaluation parameter P is less than the preset value and marking them as abnormal data sources, those data sources that are currently performing poorly can be identified in a timely manner. Conduct regular monitoring on these abnormal data sources, continuously collect their retrieved evaluation parameter P and plot a line chart, and the change trend can be visually observed. When the trend of the line chart is in a flat state or a downward state, it means that the value of the data source may be continuously decreasing or has stabilized at a low level. Generate data source abnormal information in a timely manner, enabling relevant personnel to quickly detect potential problems of the data source and make preparations in advance.

[0095] Guarantee the stability of data quality. In the high-speed audit system, the quality of data sources directly affects the accuracy and reliability of audit results. The monitoring and processing of abnormal data sources help ensure that the data sources used by the system always maintain a high quality level. Once it is found that the trend of abnormal data sources is not good, measures can be taken in a timely manner, such as replacing the data source or optimizing the data source, to avoid deviations in audit results caused by the decline in data source quality, and guarantee the stability of the data quality of the entire system.

[0096] For those abnormal data sources with low evaluation parameters and poor trends, it may no longer be worth investing too many system resources in data retrieval and processing. By discovering and warning these abnormal situations in a timely manner, the system can reasonably adjust the resource allocation strategy, concentrating resources on more valuable and stable data sources. This not only improves the resource utilization efficiency but also reduces unnecessary resource waste, enabling the system to operate more efficiently.

[0097] The generation process of the warning information is as follows:

[0098] First, construct an audit model, and the audit model includes a path model and a cost model;

[0099] After that, collect vehicle information and compare the vehicle information with the path model and the cost model;

[0100] When any comparison is abnormal, a warning message is generated.

[0101] Furthermore, the construction process of the path model is as follows: Extract all the entrance stations and exit stations from the data after ETL processing;

[0102] After that, use a preset algorithm to calculate the shortest path from all the entrance stations to the exit stations, that is, obtain the path model. The preset algorithm includes Dijkstra algorithm or A* algorithm;

[0103] The construction process of the cost model is as follows;

[0104] Extract the unit mileage charging standards and the shortest paths of different vehicle types, mark the unit mileage charging standards of different vehicle types as Fi, and mark the shortest path as G, where i is the number of vehicle types;

[0105] Through the formula G * Fi = Gf, that is, obtain the cost model Gf.

[0106] Extract the vehicle information, which includes vehicle type, vehicle entrance station, vehicle exit station and exit payment;

[0107] Extract the path model corresponding to the vehicle entrance station and vehicle exit station from all the path models;

[0108] After that, collect the actual driving mileage information of the vehicle through the road gantry, mark the actual driving mileage information as Z1, and mark the path model corresponding to the vehicle entrance station and vehicle exit station as Z2;

[0109] Through the formula (Z1 - Z2) * α = Zz, that is, obtain the evaluation coefficient. When the evaluation coefficient is greater than the preset value, it indicates that the comparison is abnormal and a warning message is generated. α is a correction value, 1.01 ≤ α ≤ 1.1, and α is proportional to Z2;

[0110] Extract the vehicle type, compare it with the preset data, retrieve the standard unit mileage cost corresponding to the vehicle type from the preset database, and mark it as E1;

[0111] Through the formula E1 * Z1 = Ez, obtain the standard cost Ez, mark the exit payment as E2, calculate the difference between E2 and E1, obtain the cost difference Ee. When Ee exceeds the preset range, it indicates that the comparison is abnormal and a warning message is generated;

[0112] By constructing a path model and a cost model and comparing them with the actual driving conditions of the vehicle, anomalies in the vehicle's driving path and cost payment can be detected in a timely manner. For example, when the deviation between the actual driving mileage of the vehicle and the mileage calculated by the shortest path model is large, or when the difference between the outbound payment and the standard cost exceeds the preset range, a warning message is generated, which helps to correct unreasonable charging situations in a timely manner, ensure the accuracy and fairness of charging, and safeguard the interests of both the highway operator and the vehicle owner.

[0113] Some vehicle owners may evade tolls by deliberately taking a detour or using improper means. This process can accurately monitor the driving path and cost of the vehicle, issue a warning once an anomaly is detected, effectively curb illegal acts such as toll evasion, and ensure the toll collection order of the highway and the economic benefits of the operator.

[0114] By comprehensively using the path analysis model, cost anomaly model, and vehicle behavior analysis model for auditing, multi-dimensional analysis of vehicle information is achieved. The system can automatically complete a series of operations such as data collection, model construction, and comparison analysis, and automatically generate warning messages when anomalies occur, reducing manual intervention, improving the auditing efficiency and accuracy, and reflecting the intelligent and automated level of the system.

[0115] In the construction of the path model, a preset algorithm is used to calculate the shortest path, and reasonable correction values and preset ranges are set in the evaluation coefficient and cost difference calculation, enabling the system to make accurate judgments according to different situations and improving the adaptability and reliability of the system.

[0116] The content audited by the data auditing module also includes abnormal vehicle behaviors. When abnormal vehicle behaviors are detected, warning messages are generated immediately.

[0117] The determination process of abnormal vehicle behaviors is as follows:

[0118] Extract the historical passing records of the vehicle from the data after ETL processing; the historical passing records include license plate number, vehicle type, entrance toll station, exit toll station, passing time (accurate to seconds or even milliseconds), toll amount, ETC transaction status, gantry capture time and location for each passing;

[0119] Associate and integrate the data from different data sources, including the lane toll system, gantry system, and ETC system, according to the vehicle identification, i.e., the license plate number, to form a passing file for each vehicle.

[0120] Calculate the number of times a vehicle passes through within a preset duration, set a normal passing frequency range obtained through statistical analysis of historical data. When the number of times a vehicle passes through exceeds the upper limit of this range, it is determined as abnormal and a warning message is generated. For example, under normal circumstances, the average number of times a small passenger car passes through a certain section per week is 3 - 5 times, while a certain vehicle reaches more than 15 times and there is no reasonable explanation for its operation (such as the lack of logistics vehicle filing information, etc.), then it is marked as an abnormal passing frequency.

[0121] Extract the passing intervals of the vehicle from the vehicle passing records, use time series analysis techniques to model the passing interval time of the vehicle. For example, use the ARIMA model to predict the reasonable passing time interval of the vehicle next time. If the deviation between the actual passing interval and the predicted value exceeds the preset value, for example, the predicted interval is 2 days, but it passes through again within 1 hour actually, it indicates that the vehicle's behavior does not conform to the normal passing pattern and is determined as abnormal, and a warning message is generated.

[0122] Based on the historical traffic flow data of the highway, divide the peak periods (such as the morning and evening rush hours on weekdays, the outbound and return peak hours on holidays, etc.). Count the number of times a vehicle passes through during the peak period and the total number of times it passes through from the vehicle passing records. If the ratio of the number of times a vehicle passes through during the peak period to the total number of times it passes through exceeds the preset value (for example, more than 80%, the ratio of normal vehicles may be between 30% - 50%), and through the analysis of the vehicle passing records, it is known that the vehicle is not an emergency rescue vehicle or a bus dedicated line vehicle, then it is determined as abnormal and a warning message is generated;

[0123] Analyze the vehicle passing records to identify that the vehicle frequently shuttles between toll stations or section combinations with a distance less than the preset value, that is, set a distance threshold (such as within 50 kilometers). If the vehicle shuttles between this short - distance interval multiple times (5 times or more) within the preset time period (such as within 3 consecutive days), and the fluctuation range of the toll amount each time is less than the preset value, this may conform to the characteristics of card - swapping to evade tolls or taking advantage of the short - distance billing loopholes, and it is determined as abnormal and a warning message is generated;

[0124] For different vehicle types, count their average toll level per kilometer on each section, establish a vehicle - type - section toll benchmark model. Analyze the vehicle passing records. When it is found that within the preset time period (such as one month), the driving mileage of the vehicle is greater than the preset value, and the ratio of the normal toll amount of the same vehicle type to the total toll amount of the current vehicle is less than the preset value, and the application of the free - passage policy is excluded, it is determined as abnormal and a warning message is generated;

[0125] Ensuring fair toll collection and stable revenue By monitoring various aspects such as the frequency of passage and the toll amount, it is possible to effectively identify violations such as evading tolls. For example, frequent short-distance round trips with little fluctuation in toll amount, abnormal vehicle-type-road section tolls, etc. may be manifestations of card swapping to avoid tolls or taking advantage of billing loopholes. Timely discovery of these abnormalities and generation of warning messages can avoid revenue losses for highway operators, ensure the fairness of toll collection, and maintain normal toll collection order.

[0126] Ensuring that each vehicle pays the fee according to the specified standard helps to stabilize the capital inflow of the operator and provides sufficient financial support for the construction, maintenance, and management of highways.

[0127] Optimizing resource allocation In-depth analysis of vehicle passage behavior enables the operator to understand the vehicle usage conditions at different times and sections. For example, by counting the passage proportion during peak hours, it is possible to know which vehicles overuse road resources during peak hours. The operator can, based on this information, reasonably plan resource allocation, such as adjusting toll policies, optimizing lane settings, etc., to improve the utilization efficiency of road resources.

[0128] For vehicles with extremely high passage frequencies, the operator can further investigate their operation models. If there is unreasonable resource occupancy, corresponding measures can be taken for guidance or restriction to make more rational use of resources.

[0129] Improving operation efficiency The automated inspection process of abnormal vehicle behaviors reduces the workload and time cost of manual inspection. The system can quickly process a large amount of vehicle passage data, timely discover abnormalities and issue warnings, enabling the operator to quickly respond and handle problems, thus improving the efficiency of operation management.

[0130] Timely discovery and handling of abnormal behaviors help to avoid problems such as traffic congestion and toll disputes caused by violations, ensure the smooth operation of highways, and improve the overall operation efficiency.

[0131] Identifying potential safety risks Vehicles with abnormal passage intervals may have safety hazards such as fatigue driving and illegal transportation. For example, if the actual passage interval of a vehicle deviates greatly from the predicted value and it passes multiple times in a short period, it may mean that the driver is fatigued, increasing the risk of traffic accidents. Timely discovery of these abnormal behaviors and intervention can effectively prevent the occurrence of safety accidents and ensure road traffic safety.

[0132] For vehicles that frequently pass during peak hours and are not emergency rescue or bus dedicated lines, their behaviors may affect normal traffic order and increase the likelihood of traffic congestion and accidents. By monitoring and handling such abnormal behaviors, potential safety risks can be reduced.

[0133] Strengthening emergency management capabilities. Understanding the normal traffic patterns and abnormal behavior models of vehicles helps the operator better conduct emergency management in case of emergencies. For example, in the event of natural disasters, traffic accidents and other emergencies, resources can be quickly allocated based on vehicle traffic records and abnormal situations, guiding vehicles to pass in an orderly manner and improving the emergency response ability.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. A high-speed audit system based on data ETL processing, characterized in that, Including: A data source acquisition module, which is used to collect information related to the data source; A data source analysis module, which is used to analyze the information related to the data source to obtain data source analysis information, and the data source analysis information includes passed evaluation and failed evaluation; A data retrieval module, which is used to select a data source for data retrieval from the data sources that have passed the evaluation, and perform ETL processing on the retrieved data; A data auditing module, which is used to audit the data after ETL processing to generate warning information; An information sending module, which is used to send the generated warning information to a preset receiving terminal.

2. The high-speed audit system based on data ETL processing according to claim 1, wherein: The specific content of the data source acquisition module for collecting information related to the data source is as follows: collecting the information of the data source being retrieved and the information of the data source being adopted; The specific process of analyzing the information related to the data source is as follows: extracting the information related to the data source, and obtaining the information of the data source being retrieved and the information of the data source being adopted from the information related to the data source; The information of the data source being retrieved includes the number of times each source data source is retrieved and the time point of each retrieval; The information of the data source being adopted is the number of times the data source is finally adopted; Processing the number of times each data source is retrieved, the time point of each retrieval, and the number of times the data source is finally adopted to obtain a retrieval evaluation parameter. When the retrieval evaluation parameter is greater than a preset value, it means that its evaluation has passed.

3. The high-speed auditing system based on data ETL processing according to claim 2, characterized in that: The process of obtaining the retrieval evaluation parameter is as follows: Mark the number of times each data source is retrieved as C, that is, the total number of times the data source is retrieved during the operation of the system; Mark the time point of each retrieval of the data source as T=(t1, t2, t3... tC); Mark the number of times the data source is finally adopted as A; For data sources with the retrieval count C > 1, calculate the time interval Δt between two adjacent retrievals i = t i+1 - t i , where i = 1, 2,..., C - 1; After calculating all the Δt i calculate the mean value to obtain the average retrieval interval U; Extract the first retrieval time and the last retrieval time of the number of times C of the data source being retrieved, calculate the difference between the last retrieval time and the first retrieval time, and obtain the total retrieval time Tg; Calculate the ratio of the total retrieval Tg to C to obtain the number of retrievals per unit time N; Then calculate the ratio of the number of times A that the data source is finally adopted to the number of times C that the data source is retrieved to obtain the adoption ratio Ac; Finally, through the formula (U + N + Ac) / 3 = P, that is, the retrieval evaluation parameter P is obtained.

4. A high-speed audit system based on data ETL processing according to claim 2, characterized in that: The process of the data retrieval module selecting a data source from the data sources that have passed the evaluation is as follows: extracting the retrieval evaluation parameter P of all data sources, and setting the data sources corresponding to the two largest retrieval evaluation parameters P as the selected data sources.

5. The high-speed audit system based on data ETL processing according to claim 2, wherein: Extract the data sources whose retrieval evaluation parameter P is less than the preset value, mark them as abnormal data sources, and conduct regular monitoring on the abnormal data sources to monitor the change of the subsequent retrieval evaluation parameter P of the abnormal data sources, that is, regularly collect the retrieval evaluation parameter P of the abnormal data sources, draw a line chart of the retrieval evaluation parameter P of the abnormal data sources in chronological order of collection time, and then analyze the trend of the line chart. When the trend of the line chart is in a flat state or a downward state, data source abnormal information is generated.

6. The high-speed auditing system based on data ETL processing according to claim 1, characterized in that: The process of generating the warning information is as follows: First, construct an auditing model, and the auditing model includes a path model and a cost model; Then collect vehicle information, and compare the vehicle information with the path model and the cost model; When any comparison is abnormal, a warning message is generated.

7. The high-speed auditing system based on data ETL processing according to claim 6, wherein: The construction process of the path model is as follows: Extract all the entry stations and exit stations from the data after ETL processing; After that, use a preset algorithm to calculate the shortest path from all entry stations to exit stations, that is, obtain the path model; The construction process of the cost model is as follows: Extract the unit mileage charging standards and the shortest paths of different vehicle types. Mark the unit mileage charging standards of different vehicle types as Fi, and the shortest path as G, where i is the number of vehicle types; Through the formula G * Fi = Gf, the cost model Gf is obtained.

8. The high-speed auditing system based on data ETL processing according to claim 6, characterized in that: Extract vehicle information, which includes vehicle type, vehicle entry station, vehicle exit station, and payment at exit; Extract the path model corresponding to the vehicle entry station and vehicle exit station from all the path models; After that, collect the actual driving mileage information of the vehicle through the road gantry. Mark the actual driving mileage information as Z1, and the path model corresponding to the vehicle entry station and vehicle exit station as Z2; Through the formula (Z1 - Z2) * α = Zz, the evaluation coefficient is obtained. When the evaluation coefficient is greater than the preset value, it indicates that the comparison is abnormal and a warning message is generated. α is a correction value, 1.01 ≤ α ≤ 1.1, and α is proportional to Z2; Extract the vehicle type and compare it with the preset data. Retrieve the standard unit mileage cost corresponding to the vehicle type from the preset database and mark it as E1; Through the formula E1 * Z1 = Ez, the standard cost Ez is obtained. Mark the payment at exit as E2, calculate the difference between E2 and E1 to obtain the cost difference Ee. When Ee exceeds the preset range, it indicates that the comparison is abnormal and a warning message is generated.

9. The high-speed auditing system based on data ETL processing according to claim 1, wherein: The content audited by the data auditing module also includes abnormal vehicle behaviors. When abnormal vehicle behaviors are found, a warning message is generated.

10. A high-speed audit system based on data ETL processing according to claim 1, characterized in that: The determination process of abnormal vehicle behaviors is as follows: Associate and integrate the data from different data sources, including the lane toll system, gantry system, and ETC system, according to the vehicle identification, that is, the license plate number, to form a traffic record for each vehicle; Extract the number of vehicle passages from the vehicle traffic record, calculate the number of vehicle passages within a preset time period, and set a normal passage frequency range obtained through statistical analysis of historical data. When the number of vehicle passages exceeds the upper limit of this range, it is determined as abnormal and a warning message is generated; Extract the passage interval of the vehicle from the vehicle traffic record, use time series analysis technology to model the passage interval time of the vehicle, predict the next reasonable passage time interval of the vehicle. If the deviation between the actual passage interval and the predicted value exceeds the preset value, it is determined as abnormal and a warning message is generated; Based on the historical traffic flow data of the highway, divide the peak periods. Count the number of vehicle passages and the total number of vehicle passages of the vehicle during the peak periods from the vehicle traffic record. If the ratio of the number of vehicle passages of the vehicle during the peak periods to the total number of vehicle passages exceeds the preset value, and after analyzing the vehicle traffic record, it is found that the vehicle is not an emergency rescue vehicle or a bus special line vehicle, it is determined as abnormal and a warning message is generated; Analyze the vehicle passing records to identify toll stations or combinations of road sections where vehicles frequently travel back and forth within a distance less than a preset value. That is, set a distance threshold. If a vehicle travels back and forth within this short-distance range multiple times within a preset time period and the fluctuation range of the toll amount each time is less than the preset value, it is determined as abnormal and a warning message is generated; For different vehicle types, calculate the average toll per kilometer on each road section, establish a vehicle type-road section toll benchmark model, and analyze the vehicle passing records. When it is found that within a preset time period, the driving mileage of a vehicle is greater than the preset value, the ratio of the normal toll amount of the same vehicle type to the total toll amount of the current vehicle is less than the preset value, and the application of the free passage policy is excluded, it is determined as abnormal and a warning message is generated.