Highway toll evasion intelligent auditing method based on multi-source data fusion
By integrating multi-source data and using automated auditing methods, the problems of low efficiency and high error rate of traditional auditing methods have been solved. This has enabled accurate identification and dynamic adaptation of toll evasion on highways, improved auditing accuracy and work efficiency, and ensured fairness in toll collection and the legitimate rights and interests of operators.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG EXPRESSWAY LINYI DEV CO LTD
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-15
AI Technical Summary
Traditional highway auditing methods are inefficient and have a high error rate, making it difficult to identify diverse toll evasion behaviors. Existing data fusion technologies have failed to fully explore the correlation characteristics of multi-source data, resulting in insufficient auditing accuracy and efficiency.
By adopting a multi-source data fusion strategy, a fully automated auditing method is constructed through distributed data acquisition, data preprocessing, feature extraction, and a multi-dimensional evasion identification model, combined with data correlation calculation and path quantification early warning, to achieve accurate identification and dynamic adaptation of diverse evasion behaviors.
Significantly reduce missed inspections and misjudgments, improve audit reliability and work efficiency, reduce toll revenue loss, ensure fair toll collection and legitimate operational rights, and achieve standardized and efficient audit work.
Smart Images

Figure CN122050134A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of traffic management technology, and in particular to an intelligent auditing method for highway toll evasion based on multi-source data fusion. Background Technology
[0002] With the continuous expansion of the expressway network and the rapid growth of motor vehicle ownership, expressway traffic volume is experiencing explosive growth, and the auditing pressure on the toll management system is becoming increasingly prominent. As a key link in ensuring toll fairness and protecting the legitimate rights and interests of operators, the efficiency and accuracy of expressway toll auditing directly affect the overall effectiveness of expressway operation and management.
[0003] Traditional highway auditing primarily relies on manual verification, requiring staff to sift through massive amounts of gantry transaction data, toll station traffic flow data, and various special vehicle data to identify suspected toll evasion. However, this manual auditing method has significant limitations: firstly, the millions of multi-source data points generated daily result in extremely low efficiency for manual screening, making comprehensive coverage difficult, and the long hours of intense work can easily lead to fatigue-induced missed detections; secondly, toll evasion behaviors are diverse and concealed, including various types such as "falsely reporting entry points," "obstructing passage media," and "large vehicles using small labels," with significant differences in behavioral characteristics between different toll evasion methods. Relying solely on staff experience for judgment is prone to misjudgment and fails to develop targeted auditing strategies.
[0004] While some existing technologies attempt to incorporate data processing techniques to assist in auditing, they often suffer from insufficient data fusion, limited feature extraction, and inadequate model adaptability. Most solutions only construct identification rules for single types of toll evasion behavior, failing to fully explore the correlation features between multi-source data. This results in limited ability to identify complex toll evasion behaviors, a lack of dynamic optimization mechanisms, and difficulty in adapting to the ever-evolving needs of toll evasion methods. Audit accuracy and efficiency still need improvement. Therefore, this paper proposes an intelligent toll evasion auditing method for highways based on multi-source data fusion to address the aforementioned problems. Summary of the Invention
[0005] The purpose of this invention is to solve the problems in the background art and to propose an intelligent auditing method for highway toll evasion based on multi-source data fusion.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: an intelligent auditing method for highway toll evasion based on multi-source data fusion, comprising the following steps: S1: Multi-source data acquisition: Acquire highway gantry transaction data, toll station traffic flow data, and related data of specific types of vehicles to construct the original dataset; S2: Data preprocessing: Based on data cleaning rules, outlier removal, missing value completion, and format standardization are performed on the original dataset to obtain standardized data; S3: Feature Extraction: Extract vehicle identity features, travel route features, billing-related features, and behavior pattern features from standardized data to construct a vehicle audit feature vector; S4: Toll evasion behavior identification: Input the vehicle audit feature vector into the preset multi-dimensional toll evasion identification model, and calculate the vehicle toll evasion probability value through the model. When the toll evasion probability value is greater than the preset threshold, it is judged as a suspected toll evasion vehicle. The multi-dimensional toll evasion identification model is trained based on machine learning algorithm and can adapt to the identification needs of various toll evasion types such as "blocking the passage medium", "large vehicle with small label" and "false entry". S5: Audit Result Output: Generates an audit report for suspected toll evasion vehicles, including feature matching details and the basis for calculating the probability of toll evasion.
[0007] In the above-mentioned intelligent audit method for highway toll evasion based on multi-source data fusion, in step S1, the collection of multi-source data is achieved through a distributed data collection architecture, and the data transmission process adopts an encryption protocol to meet the requirements of real-time performance and security; the specific type of vehicle-related data includes green channel vehicle verification data and special operation vehicle registration data.
[0008] In the above-mentioned intelligent auditing method for highway toll evasion based on multi-source data fusion, in step S2, the data cleaning rules are constructed using the following formula: In the formula, For the first in the original dataset The first data item Each attribute value These are the standardized attribute values after cleaning. For the first The mean of each attribute, For the first The standard deviation of each attribute The outlier determination coefficient is used when... When this happens, the attribute value is determined to be an outlier and is removed.
[0009] In the aforementioned intelligent auditing method for highway toll evasion based on multi-source data fusion, the toll-related features in S3 include the vehicle's rated load capacity, mileage, toll rate coefficient for toll period, and gantry toll node matching degree. During feature extraction, a feature weight allocation formula is used to weight each feature. In the formula, For the first The weight values of each feature, For the first Information gain value of each feature The formula represents the total number of features, and it assigns higher weights to features with high discriminative power.
[0010] In the above-mentioned intelligent auditing method for highway toll evasion based on multi-source data fusion, the training process of the multi-dimensional toll evasion identification model in step S4 includes: Construct a training dataset containing normal passage samples and various toll evasion samples, and normalize the sample data; The base model is constructed using the gradient boosting algorithm, and the model loss function is minimized through iterative training. In the formula, The model loss function, For the true labels of the samples, These are the model's predicted values. The regularization coefficient is . This is a penalty term for model complexity. A branch for identifying fare evasion types is introduced, and a dedicated identification sub-model is set up for different fare evasion types. A comprehensive fare evasion probability value is output through a model fusion strategy.
[0011] In the above-mentioned intelligent auditing method for highway toll evasion based on multi-source data fusion, in step S4, the preset threshold is determined through statistical learning methods, and a threshold optimization function is constructed based on historical auditing data. In the formula, To be the optimal threshold, Threshold The corresponding audit accuracy rate, Threshold The corresponding audit recall rate, This is the balancing coefficient, which is used to achieve the optimal balance between precision and recall.
[0012] In the aforementioned intelligent toll evasion audit method for highways based on multi-source data fusion, the method further includes a model update step: periodically collecting new audit data as incremental training samples, and updating the parameters of the multi-dimensional toll evasion identification model through an incremental learning formula. In the formula, For the updated model parameters, These are the model parameters before the update. For learning rate, The gradient of the loss function is calculated based on incremental samples, enabling the model to adaptively adjust to changes in toll evasion behavior.
[0013] In the aforementioned intelligent auditing method for highway toll evasion based on multi-source data fusion, the multi-source data fusion process uses a data correlation degree calculation formula to assign fusion weights to data from different sources: In the formula, For the first The fusion weight of each data source For the first The correlation between each data source and the audit target. The total number of data sources is set to ensure that highly relevant data sources play a leading role in the integration process.
[0014] In the above-mentioned intelligent auditing method for highway toll evasion based on multi-source data fusion, the travel path characteristics are quantified using a path matching degree calculation formula: In the formula, For path matching degree, This represents the actual length of the vehicle's travel route. To find the shortest path length based on the entrance and exit, The path deviation coefficient is when When the path is below the preset matching threshold, an abnormal path warning is triggered.
[0015] In the aforementioned intelligent auditing method for highway toll evasion based on multi-source data fusion, the method further includes a toll evasion behavior tracing step: for vehicles suspected of toll evasion, key evidence data is traced using a data tracing formula. In the formula, For the collection of evidence data for tracing the source, For a subset of gantry transaction data, This is a subset of toll station traffic data. As a data correlation threshold, data with a higher correlation to the determination of fare evasion are filtered out. The data serves as core evidence.
[0016] Compared with existing technologies, the advantages of this invention are: 1. This invention achieves accurate identification of diverse toll evasion behaviors through a multi-source data fusion strategy and branched model design, resulting in a significant reduction in missed detections and false judgments, and an improvement in audit reliability. It relies on the data correlation calculation formula to perform weighted fusion of multi-source data such as gantry transactions and toll station traffic flow, fully exploring the potential correlation features between data. It constructs exclusive identification sub-models for different toll evasion types such as "false entry" and "large vehicle with small label", and combines loss function optimization training and probability quantification judgment mechanism to make toll evasion behavior identification more targeted, effectively reducing the risk of missed detections and false judgments, and providing scientific and traceable quantitative basis for audit decisions.
[0017] 2. This invention achieves a leapfrog improvement in audit efficiency through a fully automated processing and dynamic update mechanism, reducing the burden on manual labor and adapting to the evolution of evasion behavior. By using standardized data preprocessing algorithms, automatic feature extraction models, and threshold optimization judgment processes, it replaces the traditional manual screening of massive amounts of data and experience-based judgment, realizing fully automated operation from data collection to audit report output. Through incremental learning formulas, the model parameters are updated regularly, enabling the model to adapt to the changing trends of evasion behavior, shortening the audit cycle, reducing the workload of auditors, and continuously ensuring the efficiency and adaptability of audit work.
[0018] 3. This invention achieves standardized toll management and protection of rights through path quantification early warning and complete evidence chain tracing design, resulting in reduced toll revenue loss and maintenance of toll fairness. It uses a path matching degree calculation formula to quantify and analyze vehicle travel paths, triggering abnormal path warnings in a timely manner; and uses a data tracing formula to screen core evidence data highly correlated with toll evasion determination, forming a complete and traceable evidence chain, making audit work rigorous and compliant, effectively curbing various toll evasion behaviors, reducing toll revenue loss, and protecting the legitimate rights and interests of highway operators. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating an intelligent audit method for highway toll evasion based on multi-source data fusion proposed in this invention. Detailed Implementation
[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Reference Figure 1 As shown, a smart auditing method for highway toll evasion based on multi-source data fusion is characterized by the following steps: S1: Multi-source data acquisition: A distributed data acquisition architecture is adopted, deploying multiple acquisition nodes to connect to the highway gantry tolling system, toll station management system, and special vehicle management system respectively, to achieve comprehensive acquisition of multi-source data. The acquired data specifically includes: gantry transaction data, toll station traffic flow data, and specific type vehicle association data. Gantry transaction data specifically includes the timestamp of the vehicle passing through the gantry, the unique identifier of the gantry, and the vehicle identification information. Toll station traffic flow data specifically includes the entrance station identifier, exit station identifier, traffic medium type, and payment record. Specific type vehicle association data specifically includes green channel vehicle verification information, special operation vehicle registration parameters, and verification attributes. The data transmission phase employs an encrypted communication protocol to ensure the security and integrity of data transmission. Simultaneously, a real-time synchronization mechanism integrates the data collected by each node into the original data set in real time, meeting the timeliness requirements of auditing work. After collection, the original data undergoes an initial format check to ensure the integrity of data fields and logical coherence, laying the foundation for subsequent preprocessing stages. S2: Data Preprocessing: Data preprocessing is based on standardization rules and mathematical formulas. The specific process is as follows: S21: Outlier Removal: First, calculate the mean of each attribute in the original dataset. and standard deviation Set outlier determination coefficient , Values are determined based on the characteristics of highway audit data; for the first... The first data item Attribute values Through formula Anomaly detection is performed. If the condition is met, the attribute value is determined to be an anomaly, and the data is removed to prevent abnormal data from interfering with subsequent analysis. S22: Missing value completion: For missing attribute values in the dataset, based on the distribution characteristics of data with the same attribute, interpolation or completion algorithms based on attribute correlation are used to fill in the missing values to ensure data integrity. S23: Format Standardization: Unify the field format, units of measurement, and encoding rules of all data, and convert heterogeneous data output from different systems into standardized data. The conversion logic follows a formula. The standardized attribute values Within a unified unit of measurement, data adaptability is improved; S3: Feature Extraction: Extract four core features from standardized data, namely vehicle identity features (vehicle identification code, license plate code, vehicle type identifier), travel route features (gantry passage sequence, road segment travel time, path node matching status), billing-related features (vehicle rated load capacity, travel mileage, travel time toll rate coefficient, gantry tolling node matching degree), and behavioral pattern features (historical travel frequency, payment habits, route selection preferences). Using the feature weight allocation formula Each feature is weighted; among them, For the first The weight values of each feature, The information gain value for this feature is obtained by calculating the degree of correlation between the feature and the toll evasion behavior. The total number of features is calculated using this formula, which gives greater weight to features that contribute more to the identification of toll evasion, thereby constructing a vehicle audit feature vector with strong discriminative ability. S4: Identification of Fee Evasion Behavior S41: Training of a multi-dimensional toll evasion detection model: S411: Training Dataset Construction: Collect a large amount of historical passage data, filter out normal passage samples and various toll evasion samples such as "false entry", "blocking passage medium" and "large vehicle with small label" to form a training dataset, and eliminate the difference in the scale between feature dimensions through normalization processing. S412: Basic Model Construction: The basic model is constructed using the gradient boosting algorithm, and the model loss function is minimized through iterative training. ;in, The model loss function, The actual labels for the samples are: 0 for normal passage and 1 for fare evasion. These are the model's predicted values. The regularization coefficient is . This is a penalty term for model complexity. The loss function is used to balance the model's fitting ability and generalization ability to avoid overfitting. S413: Multi-branch model fusion: To address the differences in behavioral characteristics of different types of toll evasion, a dedicated identification sub-model is built on the basic model. Each sub-model focuses on the feature identification of one type of toll evasion. Through a model fusion strategy, the probability values of a single toll evasion type output by each sub-model are integrated into a comprehensive toll evasion probability value, thereby improving the model's adaptability to various toll evasion behaviors. S42: Fee Evasion Detection Process: The vehicle's audit feature vector is input into the trained multi-dimensional toll evasion detection model. The model calculates the vehicle's comprehensive toll evasion probability value using its built-in algorithm; a preset threshold is then applied. Through threshold optimization function Confirmed, among which Threshold The corresponding audit accuracy rate, Threshold The corresponding audit recall rate, For balance coefficient, Adjusted according to actual audit needs, this function achieves an optimal balance between accuracy and recall rate; when the overall probability of vehicle toll evasion exceeds a preset threshold... At that time, the vehicle was determined to be a suspected toll evader, and the details of feature matching and probability calculation were recorded; S5: Audit Result Output: The system automatically generates a standardized audit report for suspected toll evasion vehicles. The report includes basic vehicle information, details of multi-source data, feature extraction results, the basis for calculating the probability of toll evasion, including the formula application process, and feature matching details. The audit report is pushed to the toll audit center management platform through the system interface, supporting staff to view, retrieve, and conduct subsequent verification online, providing clear and traceable reference for audit work. S6: Model Update: To adapt to the dynamic evolution of toll evasion behavior, a regular model update mechanism is set up; new audit data, including verified normal passage data and toll evasion data, are collected periodically as incremental training samples, and the model is learned through an incremental learning formula. Update the model parameters; among them, For the updated model parameters, These are the model parameters before the update. The learning rate controls the step size for updating parameters. The gradient of the loss function is calculated based on incremental samples; through continuous parameter updates, the model can adapt to new evasion behavior features in real time and maintain stable recognition performance. S7: Data Fusion and Auxiliary Processing S71: Multi-source data fusion: using the data correlation degree calculation formula Assign fusion weights to each data source, where For the first The fusion weight of each data source The correlation between this data source and the audit objectives is determined by analyzing the contribution of the data to the identification of toll evasion. This represents the total number of data sources. This formula ensures that highly correlated data sources dominate the fusion process, guaranteeing that the fused data accurately reflects the correlation between vehicle traffic status and toll evasion behavior. S72: Path Feature Quantization: Calculated using the path matching degree formula The characteristics of the travel path are quantified, among which For path matching degree, This represents the actual length of the vehicle's travel route. To find the shortest path length based on the entrance and exit, This is the path deviation coefficient. Based on the highway network structure; when the calculated path matching degree is set... When the matching degree is below the preset threshold, the system automatically triggers a path anomaly warning and marks the vehicle as a key audit target; S73: Tracing the source of toll evasion: For vehicles suspected of evading tolls, trace the source using data tracing formulas. Screening core evidence data; among which, For the collection of evidence data for tracing the source, For a subset of gantry transaction data, This is a subset of toll station traffic data. The threshold for data correlation is used to filter out data with a higher correlation to the determination of fare evasion. The data serves as core evidence, forming a complete chain of evidence that provides strong support for subsequent audits, investigations, and handling of violations.
[0022] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A smart auditing method for highway toll evasion based on multi-source data fusion, characterized in that, Includes the following steps: S1: Multi-source data acquisition: Acquire highway gantry transaction data, toll station traffic flow data, and related data of specific types of vehicles to construct the original dataset; S2: Data preprocessing: Based on data cleaning rules, outlier removal, missing value completion, and format standardization are performed on the original dataset to obtain standardized data; S3: Feature Extraction: Extract vehicle identity features, travel route features, billing-related features, and behavior pattern features from standardized data to construct a vehicle audit feature vector; S4: Toll evasion behavior identification: Input the vehicle audit feature vector into the preset multi-dimensional toll evasion identification model, and calculate the vehicle toll evasion probability value through the model. When the toll evasion probability value is greater than the preset threshold, it is judged as a suspected toll evasion vehicle. The multi-dimensional toll evasion identification model is trained based on machine learning algorithm and can adapt to the identification needs of various toll evasion types such as "blocking the passage medium", "large vehicle with small label" and "falsely reporting the entrance". S5: Audit Result Output: Generates an audit report for suspected toll evasion vehicles, including feature matching details and the basis for calculating the probability of toll evasion.
2. The method according to claim 1, characterized in that, In S1, the collection of multi-source data is achieved through a distributed data collection architecture, and the data transmission process adopts an encryption protocol to meet the requirements of real-time performance and security. The specific type of vehicle-related data includes green channel vehicle verification data and special operation vehicle registration data.
3. The method according to claim 1, characterized in that, In step S2, the data cleaning rules are constructed using the following formula: In the formula, For the first in the original dataset The first data item Each attribute value These are the standardized attribute values after cleaning. For the first The mean of each attribute, For the first The standard deviation of each attribute The outlier determination coefficient is used when... When this happens, the attribute value is determined to be an outlier and is removed.
4. The method according to claim 1, characterized in that, In S3, the billing-related features include the vehicle's approved load capacity, travel mileage, toll rate coefficient for travel period, and gantry billing node matching degree. During the feature extraction process, a feature weight allocation formula is used to weight each feature: In the formula, For the first The weight values of each feature, For the first Information gain value of each feature The formula represents the total number of features, and it assigns higher weights to features with high discriminative power.
5. The method according to claim 1, characterized in that, In S4, the training process of the multi-dimensional fare evasion detection model includes: Construct a training dataset containing normal passage samples and various toll evasion samples, and normalize the sample data; The base model is constructed using the gradient boosting algorithm, and the model loss function is minimized through iterative training. In the formula, The model loss function, For the true labels of the samples, These are the model's predicted values. The regularization coefficient is . This is a penalty term for model complexity. A branch for identifying fare evasion types is introduced, and a dedicated identification sub-model is set up for different fare evasion types. A comprehensive fare evasion probability value is output through a model fusion strategy.
6. The method according to claim 1, characterized in that, In step S4, the preset threshold is determined through a statistical learning method, and a threshold optimization function is constructed based on historical audit data. In the formula, To be the optimal threshold, Threshold The corresponding audit accuracy rate, Threshold The corresponding audit recall rate, This is the balancing coefficient, which is used to achieve the optimal balance between precision and recall.
7. The method according to claim 1, characterized in that, The method also includes a model update step: periodically collecting new audit data as incremental training samples, and updating the parameters of the multi-dimensional evasion identification model through an incremental learning formula. In the formula, For the updated model parameters, These are the model parameters before the update. For learning rate, The gradient of the loss function is calculated based on incremental samples, enabling the model to adaptively adjust to changes in toll evasion behavior.
8. The method according to claim 1, characterized in that, The multi-source data fusion process uses a data correlation calculation formula to assign fusion weights to data from different sources: In the formula, For the first The fusion weight of each data source For the first The correlation between each data source and the audit target. The total number of data sources is set to ensure that highly relevant data sources play a leading role in the integration process.
9. The method according to claim 1, characterized in that, The characteristics of the travel path are quantified using a path matching degree calculation formula: In the formula, For path matching degree, This represents the actual length of the vehicle's travel route. To find the shortest path length based on the entrance and exit, The path deviation coefficient is when When the path is below the preset matching threshold, an abnormal path warning is triggered.
10. The method according to claim 1, characterized in that, The method also includes a step for tracing the source of toll evasion: for vehicles suspected of toll evasion, key evidence data is traced using a data tracing formula. In the formula, For the collection of evidence data for tracing the source, For a subset of gantry transaction data, This is a subset of toll station traffic data. As a data correlation threshold, data with a higher correlation to the determination of fare evasion are filtered out. The data serves as core evidence.