Network charging data anomaly detection and correction method based on machine learning model

By combining multi-granularity weighted isolation forests and graph convolutional neural networks, the problem of identifying and correcting complex abnormal data in the networked toll collection system is solved, efficient and automated anomaly detection and correction is achieved, and the intelligence and operational efficiency of the system are improved.

CN120611216APending Publication Date: 2025-09-09SHANXI TRAFFIC CONTROL DIGITAL TRAFFIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510709049.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies have difficulty effectively identifying and correcting complex and hidden abnormal data in networked toll collection systems, resulting in charging deviations, increased user complaints and decreased operational efficiency. In addition, existing methods rely heavily on manual resources, have a high misjudgment rate, and lack self-learning and generalization capabilities.

Method used

A multi-granularity weighted isolation forest model and graph convolutional neural network are used to construct a multi-granularity weighted isolation forest model for anomaly scoring, map it into a directed heterogeneous graph, use the graph convolutional neural network for graph embedding training, and combine field-level error calculation and credibility judgment for automatic correction.

Benefits of technology

It improves the accuracy and robustness of anomaly detection, enables efficient and automated repair, reduces the cost of manual participation, and improves the intelligence level and operational efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611216A_ABST
    Figure CN120611216A_ABST
Patent Text Reader

Abstract

The invention discloses a machine learning model-based networking charging data anomaly detection and correction method. The method comprises the following steps of S1, collecting networking charging data and completing preprocessing; s2, constructing a multi-granularity weighted isolated forest model, generating a plurality of isolated trees, and calculating a weighted anomaly score; s3, setting an abnormal threshold value by adopting a dynamic quantile strategy based on the abnormal score set, and screening abnormal data; s4, converting the abnormal data into a directed heterogeneous graph, and establishing a field co-occurrence structural relationship; s5, based on the directed heterogeneous graph, performing embedding training on the graph structure by using the graph convolutional neural network to generate a prediction correction value; s6, performing field-level error calculation on the predicted correction value and the abnormal data, performing correction in combination with a credibility judgment rule, and generating a corrected data set; and S7, feeding back the corrected data set to the networking charging settlement center, and updating the original record. According to the invention, charging data abnormity intelligent detection and automatic correction are realized, and accuracy and processing efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data processing and data error correction, and in particular to a method for detecting and correcting anomalies in online charging data based on a machine learning model. Background Art

[0002] Amid the rapid development of the modern transportation industry, networked toll collection systems, as a crucial component of intelligent transportation, have been widely adopted in traffic management and billing for various types of roads, including highways and urban expressways. These systems collect and process dynamic vehicle information to automatically calculate and settle tolls for different roads and travel time periods. These systems typically involve multiple key processes, including vehicle identification, route tracking, toll calculation, data upload, and settlement. They rely on the accurate collection and stable transmission of data from multiple sources. However, due to various factors, such as equipment failure, environmental interference, communication delays, and manual data entry errors, the traffic data generated by these systems often contains various types of anomalies, including abnormal toll amounts, inconsistent time points, disjointed routes, license plate recognition errors, and missing or duplicate data. If these anomalies are not promptly detected and corrected, they can lead to toll discrepancies, increased user complaints, decreased operational efficiency, and even compromise the accuracy and reliability of the entire transportation settlement system.

[0003] In existing technology, anomaly detection and correction for online toll collection data often uses a rule-based approach. This involves business personnel presetting various threshold rules and anomaly type discrimination logic, and then identifying anomalous data by traversing and screening the data set. This approach offers a certain degree of operability and controllability, capable of handling known anomalies. However, with the exponential growth of data volumes in online toll collection systems and the increasing complexity and hidden nature of anomaly types, relying solely on fixed rules is no longer sufficient to address new anomaly patterns in dynamic environments. Furthermore, rule-based approaches are costly to maintain, requiring regular manual updates to the rule set to adapt to system changes. Improper rule configuration can also lead to misjudgments and missed detections. Furthermore, rule-based approaches typically lack self-learning and generalization capabilities, making it difficult to accurately identify complex anomalies with strong implicit correlations or across multiple fields, severely impacting the system's intelligence and response efficiency.

[0004] In recent years, with the continuous development of artificial intelligence and big data technologies, anomaly detection methods based on machine learning have gradually become the mainstream approach for solving anomaly problems in large-scale, complex data structures. In general scenarios, the Isolation Forest (Isolation Forest) model, an unsupervised learning model, constructs multiple random partitioning trees to model the degree of isolation of data in a high-dimensional feature space. It has the advantages of processing large-scale data, high computational efficiency, and the lack of the need for labeled data. It has been widely used in fields such as financial risk control, network security, and equipment monitoring. However, directly applying the Isolation Forest model to online toll collection scenarios has the following shortcomings: First, the traditional Isolation Forest model treats all features equally and lacks modeling of the importance of different field granularity, making it difficult to highlight the detection weight of key fields. Second, its simple structure makes it impossible to perform hierarchical analysis of the correlation patterns of different field combinations in the toll collection data. Third, the anomaly score output by the model lacks contextual interpretation in actual use, limiting the accuracy and controllability of subsequent correction steps.

[0005] On the other hand, existing technologies also have limitations in data correction. Most of the current mainstream methods are based on manual review or template correction mechanisms. For data that is judged to be abnormal, manual correction is performed through manual verification and comparison of traffic trajectories, historical records, etc. This method not only relies on a large amount of manual resources and is inefficient, but is also prone to human errors when faced with complex logic or high-frequency abnormal scenarios. At the same time, manual review can usually only handle explicit anomalies, and it is difficult to accurately correct implicit errors caused by abnormal field combinations, disordered time logic, or inconsistent cross-site paths. In addition, decisions made during the correction process often lack contextual information support, and it is impossible to understand the dependencies between fields from a global perspective, resulting in insufficient accuracy of the correction results.

[0006] To address the above issues, some studies have begun to introduce graph neural networks to model the relationships between structured data. By building a graph structure, the model can learn the implicit dependencies and structural features between fields. However, existing solutions currently focus on entity recognition and node classification tasks in a single scenario, and have not yet formed a complete solution path that can be applied to the correction of toll data anomalies. Especially in the field of networked toll collection, toll data naturally has graph structural characteristics, such as the formation of stable structural connection relationships between vehicle, time, route, and amount fields. However, existing graph learning methods have not yet effectively combined domain business knowledge with anomaly repair logic, resulting in significant room for improvement in correction accuracy and actual application effects.

[0007] Therefore, how to provide a method for detecting and correcting anomalies in online charging data based on a machine learning model is an urgent problem that technicians in this field need to solve. Summary of the Invention

[0008] One purpose of the present invention is to propose a method for detecting and correcting anomalies in networked toll collection data based on a machine learning model. The present invention comprehensively utilizes a multi-granularity weighted isolation forest algorithm and a graph convolutional neural network model, and describes in detail the complete process of efficient anomaly identification and intelligent correction of traffic data in large-scale networked toll collection scenarios. It has the advantages of high detection accuracy, high correction efficiency and strong system adaptability.

[0009] According to an embodiment of the present invention, a method for detecting and correcting anomalies in online charging data based on a machine learning model includes the following steps:

[0010] S1. Collect online charging data and perform pre-processing;

[0011] S2. Based on the preprocessed charging data, a multi-granularity weighted isolation forest model is constructed and trained. Multiple isolated trees are randomly obtained. The average path length of each charging data item in each isolated tree is counted and the corresponding anomaly score is calculated.

[0012] S3. Use a dynamic quantile strategy to determine the abnormality threshold. If the abnormality score is greater than the abnormality threshold, the charging data is determined to be abnormal data and the abnormal data set is output;

[0013] S4. Map the abnormal data set into graph-structured data, using vehicle identification, charging time, charging location, and charging amount as node fields, and establish directed edges between nodes based on the co-occurrence relationship between the fields to form a directed heterogeneous graph;

[0014] S5. Based on the directed heterogeneous graph, a graph convolutional neural network model is constructed. Graph embedding training is performed on nodes through node aggregation and edge feature propagation mechanisms to obtain a structured graph representation of each anomaly record and output the corresponding prediction correction value based on the graph representation.

[0015] S6. Calculate the field-level error between the predicted correction value and the abnormal data, and filter the abnormal records that meet the correction conditions in combination with the credibility judgment rules, and perform a field replacement operation to generate a corrected data set;

[0016] S7. Send the corrected data set to the networked charging data settlement center to replace the original abnormal record and complete the real-time update of the charging record.

[0017] Optionally, the charging data includes vehicle identification, charging time, charging location and charging amount, and the preprocessing includes data cleaning, normalization and missing value filling.

[0018] Optionally, the S2 specifically includes:

[0019] S21, extract feature fields from the pre-processed charging data, and construct a feature vector set X = {x1, x2, ..., x n}, where x i =[v i ,t i ,l i ,a i ], is the feature vector corresponding to the i-th charging record, n is the total number of charging records, v i is the vector corresponding to the vehicle identification field in the i-th record, t i is the normalized value of the charging time field in the i-th record, l i is the code value of the charging location field in the i-th record, a i is the normalized value of the charge amount field in the i-th record;

[0020] S22, based on the feature vector set X, construct multiple feature granularity subspaces, and define the granularity set as G = {g1, g2, ..., g k}, where k is the number of granular subspaces, is the jth granularity subspace composed of a field combination, wherein the field combination consists of a vehicle identification field, a charging time field, a charging location field, and a charging amount field;

[0021] S23, in each granularity subspace g j In the random subspace sampling method, m j isolated trees, forming an isolated tree set Where T jp is the granularity subspace g j The pth isolated tree constructed under j is the granularity subspace g j the number of isolated trees;

[0022] S24. For each eigenvector x i In each isolated tree T jp Get the path length, which is recorded as h(x i ,T jp ), calculate the particle size g j The average path length under , and weight the average path length of each particle size:

[0023]

[0024] in, is the particle size g j The average path length of the next i-th record, w j is the weight coefficient of the j-th granularity subspace, is the weighted average path length of the i-th record at all granularities, satisfying w j ∈[0,1], and

[0025] S25. Calculate the anomaly score of the \(i\)-th record according to the weighted average path length as follows:

[0026]

[0027] where \(s(x i ) is the anomaly score of the \(i\)-th record, \(c(n)\) is the normalization constant, which depends on the total number \(n\) of charging records, \(H(n)\) is the \(n\)-th harmonic number, \(H(n - 1)\) is the \((n - 1)\)-th harmonic number, and \(\gamma\) is the Euler constant with a value of 0.5772.

[0028] Optionally, the specific steps of S23 include:

[0029] S231. For the feature set in the granularity subspace \(g j , randomly draw samples without replacement from the feature vector set \(X\) to form a training sample set The number of training samples is where \(n\) is the total number of charging records;

[0030] S232. During the construction process of each isolation tree \(T jp , recursively execute the following steps until the maximum depth \(d max or the number of leaf node samples is 1:

[0031] S2321. Randomly select a dimension \(f\in g\) from the feature dimensions of the current node j as the splitting dimension;

[0032] S2322. Randomly select a splitting point \(q\) within the value range of the feature dimension \(f\), and divide the samples into left and right subtrees according to \(f(x)<q\);​​​​​​​​​​​​​​​​​​​​​​​​​i ) is the anomaly score of the i-th record, and n is the total number of charging records;

[0038] S32. Sort the abnormal score set S in non-descending order according to the numerical value to obtain the sorted score sequence S′={s′1,s′2,…,s′ n}, satisfying s′1≤s′2≤…≤s′ n ;

[0039] S33. Set the target quantile level q∈(0,1), use the quantile level q as the percentage index of the position corresponding to the threshold, and calculate its corresponding quantile position index P:

[0040]

[0041] Among them, q is the target quantile level, N is the number of scoring samples, s′ i is the i-th score after sorting, log(N+1) is the regularization factor, which is used to balance the impact of sample size on the quantile position. is the sum of abnormal scores, max(S′) is the maximum value in the score sequence, min(S′) is the minimum value in the score sequence, is the floor function;

[0042] S34. According to the position index P, select the Pth element from the sorted scoring sequence as the abnormality judgment threshold, set to T q , each eigenvector x i The anomaly score s(x i ) is compared with the threshold, if s(x i )>T q , then the data is judged to be abnormal data, and the abnormal data set X that meets the conditions is output abn .

[0043] Optionally, the S4 specifically includes:

[0044] S41. Extract abnormal data set X abn ={x i ∣s(x i )>T q}, and define a directed heterogeneous graph G, where x i =[v i ,t i ,l i ,a i ], is the feature vector corresponding to the i-th charging record, n is the total number of charging records, v i is the vector corresponding to the vehicle identification field in the i-th record, t i is the normalized value of the charging time field in the i-th record, l iis the code value of the charging location field in the i-th record, a i is the normalized value of the charge amount field in the i-th record, s(x i ) is the abnormal score of the i-th record, T q is the threshold;

[0045] S42, according to each abnormal record x i =[v i ,t i ,l i ,a i ], establish directed edge relationships in the graph and group all directed edges into edge sets in For directed edges, further generate the final directed heterogeneous graph G;

[0046] S43. Constructing the adjacency matrix between nodes The definition is as follows:

[0047]

[0048] Among them, A u,v is the connection relationship from node u to node v, |V| is the total number of nodes;

[0049] S44. Construct node feature matrix Each of these lines is the feature vector of the i-th node in the graph, and d is the feature dimension;

[0050] S45. Output the constructed directed heterogeneous graph structure G, adjacency matrix A and node feature matrix X V , for subsequent graph convolutional neural network model processing and calling.

[0051] Optionally, the directed isomeric graph G=(V, E, Φ, Ψ) is obtained by a graph structure definition, where:

[0052] V is the set of all nodes in the graph, satisfying V = V v ∪V t ∪V l ∪V a ;

[0053] V v ={v i} is the vehicle identification node set;

[0054] V t ={t i} is the charging time node set;

[0055] V l ={l i} is the charging location node set;

[0056] V a ={a i} is the charging amount node set;

[0057] is a set of directed edges;

[0058] Φ:V→{Vehicle,Time,Location,Amount} is the node type mapping function;

[0059] Ψ:E→{(Vehicle→Time),(Time→Location),(Location→Amount)} is the edge type mapping function.

[0060] Optionally, the directed edge relationship is through the exception record x i =[v i ,t i ,l i ,a i ] to obtain the graph structure including:

[0061] From the vehicle identification node v i Point to charging time node t i , define the directed edge

[0062] From the charging time node t i Point to charging location node l i , define the directed edge

[0063] From charging location node l i Point to the charging amount node a i , define the directed edge

[0064] Optionally, the S5 specifically includes:

[0065] S51. On the directed heterogeneous graph G = (V, E, Φ, Ψ), combine the adjacency matrix A and the node feature matrix X V , define the l-th layer embedding output of the graph convolutional neural network as Initialize embedding to H (0) =X V ,in is the node representation matrix of the 0th layer of the graph convolutional neural network, d l is the dimension of the node embedding in the lth layer, |V| is the total number of nodes;

[0066] S52. Perform graph convolution operation and define the propagation function of the l+1th layer of the graph convolutional neural network based on the embedding result of the lth layer:

[0067]

[0068] in, represents the embedding representation of the l-th layer node, is the weight matrix of the lth layer, is the bias term, I is the unit matrix, which means the introduction of self-connection edges. for The degree matrix of σ(·) is the activation function, which takes ReLU, R is the relationship set in the heterogeneous graph, including edge types, is the adjacency matrix of relationship type r, A r The degree matrix of is the weight matrix of relationship type r in layer l, Fusion weights for heterogeneous relations;

[0069] S53, repeat the graph convolution propagation function to perform L-layer embedding propagation, and finally output the graph embedding matrix:

[0070]

[0071] where d L is the final graph embedding dimension, Z is the graph embedding matrix of all nodes;

[0072] S54. For each abnormal data record x i =[v i ,t i ,l i ,a i ], extract x from the graph embedding matrix Z i The corresponding four node embedding vectors Splice them into a structured graph representation and input them into the multi-layer perceptron prediction model to obtain the prediction correction value vector:

[0073]

[0074] Among them, ‖ represents the vector concatenation operation, Represents record x i The structured graph representation of For record x i The forecast revision of These are the predicted correction results of the vehicle identification, charging time, charging location and charging amount fields;

[0075] S55. Output the predicted correction value set of all abnormal records for subsequent field error judgment and abnormality correction.

[0076] Optionally, the S6 specifically includes:

[0077] S61. For each abnormal record x i =[v i ,t i ,l i ,a i ] and its forecast revision value Calculate the absolute error between each field and construct a field-level error vector:

[0078]

[0079] in, They are defined as follows:

[0080] The value is 1 when the vehicle identifications are different, otherwise it is 0;

[0081] is the absolute error of the charging time field;

[0082] The value is 1 when the charging location is different, otherwise it is 0;

[0083] is the absolute error of the charge amount field;

[0084] is the indicator function;

[0085] S62, set the credibility judgment rule vector θ = [θ v ,θ t ,θ l ,θ a ], where ∈{0,1} is the vehicle identification correctability condition, is the maximum allowable error of charging time, θ l ∈{0,1} is the correctability condition of the charging location, The maximum permissible error in the charge amount;

[0086] According to the field-level error vector δ i With the credibility judgment rule θ, we judge whether the abnormal data of item i meets the correction condition and define the field credibility judgment function η(x i ):

[0087]

[0088] S63, for satisfying η(x i )=1, and construct the corrected data record x′ i =[v′ i ,t′ i ,l′i ,a′ i ],in:

[0089]

[0090] S64. All records that meet the correction conditions and complete field replacement are combined into a corrected data set, and the corrected data set is output for the network charging data settlement center to update the original abnormal record.

[0091] The beneficial effects of the present invention are:

[0092] First, by constructing a multi-granularity weighted isolation forest model, this paper segments and models online charging data at multiple granular levels within the feature space. This model can fully exploit the data distribution characteristics under different field combinations and effectively improve the ability to identify complex anomaly patterns. Compared with the traditional isolation forest method, this model introduces field weights during the detection process, strengthening the focus on key fields and making anomaly scoring more reasonable and accurate, thereby significantly improving the accuracy and robustness of anomaly detection.

[0093] Secondly, this invention uses a graph convolutional neural network to model the contextual structure of abnormal data, constructing a directed heterogeneous graph of fields such as vehicle identification, charging time, location, and amount. The connection relationships between nodes are used to learn the implicit dependencies between fields. This solves the problem of the existing technology lacking field correlation analysis and difficulty in achieving accurate correction. Anomalous records are predicted and corrected through a graph embedding mechanism, effectively replacing manual comparison and static rules, achieving efficient and automated repair for complex data.

[0094] Finally, this invention incorporates a data credibility assessment mechanism and field-level error analysis into the anomaly detection and correction process, ensuring the stability and operational controllability of corrective actions, reducing the rate of false positives and the risk of incorrect corrections. This overall approach exhibits strong self-learning and adaptability, enabling it to continue functioning in the dynamic environment of networked toll collection systems. This approach enhances the intelligence and operational efficiency of toll data processing systems, and possesses significant engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0096] Figure 1 This is a flow chart of the method for detecting and correcting anomalies in online charging data based on a machine learning model proposed in the present invention. DETAILED DESCRIPTION

[0097] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0098] refer to Figure 1 , a method for detecting and correcting anomalies in online charging data based on a machine learning model, comprising the following steps:

[0099] S1. Collect online charging data and perform pre-processing;

[0100] S2. Based on the preprocessed charging data, a multi-granularity weighted isolation forest model is constructed and trained. Multiple isolated trees are randomly obtained. The average path length of each charging data item in each isolated tree is counted and the corresponding anomaly score is calculated.

[0101] S3. Use a dynamic quantile strategy to determine the abnormality threshold. If the abnormality score is greater than the abnormality threshold, the charging data is determined to be abnormal data and the abnormal data set is output;

[0102] S4. Map the abnormal data set into graph-structured data, using vehicle identification, charging time, charging location, and charging amount as node fields, and establish directed edges between nodes based on the co-occurrence relationship between the fields to form a directed heterogeneous graph;

[0103] S5. Based on the directed heterogeneous graph, a graph convolutional neural network model is constructed. Graph embedding training is performed on nodes through node aggregation and edge feature propagation mechanisms to obtain a structured graph representation of each anomaly record and output the corresponding prediction correction value based on the graph representation.

[0104] S6. Calculate the field-level error between the predicted correction value and the abnormal data, and filter the abnormal records that meet the correction conditions in combination with the credibility judgment rules, and perform a field replacement operation to generate a corrected data set;

[0105] S7. Send the corrected data set to the networked charging data settlement center to replace the original abnormal record and complete the real-time update of the charging record.

[0106] This paper proposes a method for detecting and correcting anomalies in online toll collection data based on a machine learning model. This method utilizes an integrated process encompassing data acquisition, preprocessing, anomaly detection, graph modeling, graph neural network correction, field-level error judgment, and data feedback. This method enables fully automated processing of large-scale, multi-field, high-frequency dynamic data in online toll collection scenarios. This method is highly systematic and engineering-scalable, significantly improving the intelligent level of anomaly data identification in online toll collection systems, significantly reducing manual effort, and enhancing data processing efficiency and accuracy.

[0107] In this embodiment, the charging data includes vehicle identification, charging time, charging location and charging amount, and the preprocessing includes data cleaning, normalization and missing value filling.

[0108] This paper further defines the field composition and preprocessing strategy for online toll collection data, specifically including key fields such as vehicle identification, collection time, collection location, and collection amount. Furthermore, data cleaning, normalization, and missing value filling ensure the stability and standardization of model training data. This design improves the input quality for subsequent model construction, ensures the accuracy and consistency of anomaly detection and correction, and contributes to the construction of a highly reliable intelligent data processing system.

[0109] In this embodiment, S2 specifically includes:

[0110] S21, extract feature fields from the pre-processed charging data, and construct a feature vector set X = {x1, x2, ..., x n}, where x i =[v i ,t i ,l i ,a i ], is the feature vector corresponding to the i-th charging record, n is the total number of charging records, v i is the vector corresponding to the vehicle identification field in the i-th record, t i is the normalized value of the charging time field in the i-th record, l i is the code value of the charging location field in the i-th record, a i is the normalized value of the charge amount field in the i-th record;

[0111] S22, based on the feature vector set X, construct multiple feature granularity subspaces, and define the granularity set as G = {g1, g2, ..., g k}, where k is the number of granular subspaces, is the jth granularity subspace composed of a field combination, wherein the field combination consists of a vehicle identification field, a charging time field, a charging location field, and a charging amount field;

[0112] S23, in each granularity subspace g j In the random subspace sampling method, m j isolated trees, forming an isolated tree set Where T jp is the granularity subspace g j The pth isolated tree constructed under j is the granularity subspace g j the number of isolated trees;

[0113] S24. For each eigenvector xi In each isolated tree T jp Get the path length, which is recorded as h(x i ,T jp ), calculate the particle size g j The average path length under , and weight the average path length of each particle size:

[0114]

[0115] in, is the particle size g j The average path length of the next i-th record, w j is the weight coefficient of the j-th granularity subspace, is the weighted average path length of the i-th record at all granularities, satisfying w j ∈[0,1], and

[0116] S25, according to the weighted average path length Calculate the anomaly score of the i-th record:

[0117]

[0118] Among them, s(x i ) is the anomaly score of the ith record, c(n) is the normalization constant that depends on the total number of charging records n, H(n) is the nth harmonic number, H(n-1) is the n-1th harmonic number, and γ is the Euler constant, which is 0.5772.

[0119] This paper introduces a multi-granularity weighted isolation forest model, fully considering the diverse impacts of different field combinations in charging data on anomaly detection. It constructs a multi-level, multi-subspace set of isolation trees and assigns weights based on the importance of each subspace for weighted scoring. Compared to the traditional isolation forest model, this method can more accurately identify complex anomalies caused by cross-field combination features, effectively improving the model's sensitivity and coverage for hidden anomalies.

[0120] In this embodiment, the S23 specifically includes:

[0121] S231, for the granularity subspace g j The feature set in , randomly extract samples from the feature vector set X without replacement to form a training sample set The number of training samples is Where n is the total number of charging records;

[0122] S232. In each isolated tree T jp During the construction process, the following steps are recursively performed until the maximum depth d is reached. maxOr the number of leaf node samples is 1:

[0123] S2321. Randomly select a dimension f ∈ g from the feature dimensions of the current node j as the splitting dimension;

[0124] S2322. Randomly select a splitting point q within the value range of the feature dimension f, and divide the samples into a left subtree and a right subtree according to f(x) < q;

[0125] S2323. Continue to perform the splitting operation on the left and right subtrees to construct the tree structure;

[0126] S233. Set the maximum depth Ensure that the depth of the tree structure is adapted to the scale of the sample quantity to improve the isolation efficiency;

[0127] S234. Store all the constructed isolation trees T jp in the isolation forest set F at the corresponding granularity level j for subsequent path length calculation.

[0128] The present invention further defines the construction method of the isolation tree. The isolation tree is established by adopting a sampling without replacement strategy and a recursive splitting mechanism, and the maximum tree depth is set to match the data scale, so as to improve the tree construction efficiency and the model generalization ability. This structure optimization method enhances the ability of the model to identify the sparse distribution in the high-dimensional feature space, makes the isolation path length more stable and reliable, and thus improves the distribution accuracy of the anomaly score.

[0129] In this embodiment, S3 specifically includes:

[0130] S31. Obtain the anomaly score set of all feature vectors x i , denoted as S = {s(x1), s(x2), …, s(x n i)}, where s(x i i) is the anomaly score of the i-th record, and n is the total number of toll records;

[0131] S32. Sort the anomaly score set S in non-decreasing order according to the numerical value, and obtain the sorted score sequence S′ = {s′1, s′2, …, s′ n n}, satisfying s′1 ≤ s′2 ≤ … ≤ s′ n ;

[0132] S33. Set the target quantile level q ∈ (0, 1), use the quantile level q as the percentage index corresponding to the threshold position, and calculate its corresponding quantile position index P:

[0133]

[0134] Among them, q is the target quantile level, N is the number of scoring samples, s′ i is the i-th score after sorting, log(N+1) is the regularization factor, which is used to balance the impact of sample size on the quantile position. is the sum of abnormal scores, max(S′) is the maximum value in the score sequence, min(S′) is the minimum value in the score sequence, is the floor function;

[0135] S34. According to the position index P, select the Pth element from the sorted scoring sequence as the abnormality judgment threshold, set to T q , each eigenvector x i The anomaly score s(x i ) is compared with the threshold, if s(x i )>T q , then the data is judged to be abnormal data, and the abnormal data set X that meets the conditions is output abn .

[0136] This paper proposes a method for determining anomaly thresholds based on a sorting quantile strategy. Through non-descending sorting and quantile indexing, it accurately extracts a threshold from a set of anomaly scores. This method balances data size, anomaly severity, and score distribution characteristics, avoiding the overfitting problem associated with traditional static threshold settings and enhancing the system's adaptability and dynamic adjustment capabilities.

[0137] In this embodiment, the S4 specifically includes:

[0138] S41. Extract abnormal data set X abn ={x i ∣s(x i )>T q}, and define a directed heterogeneous graph G, where x i =[v i ,t i ,l i ,a i ], is the feature vector corresponding to the i-th charging record, n is the total number of charging records, v i is the vector corresponding to the vehicle identification field in the i-th record, t i is the normalized value of the charging time field in the i-th record, l i is the code value of the charging location field in the i-th record, a i is the normalized value of the charge amount field in the i-th record, s(x i ) is the abnormal score of the i-th record, T q is the threshold;

[0139] S42, according to each abnormal record x i=[v i ,t i ,l i ,a i ], establish directed edge relationships in the graph and group all directed edges into edge sets in For directed edges, further generate the final directed heterogeneous graph G;

[0140] S43. Constructing the adjacency matrix between nodes The definition is as follows:

[0141]

[0142] Among them, A u,v is the connection relationship from node u to node v, |V| is the total number of nodes;

[0143] S44. Construct node feature matrix Each of these lines is the feature vector of the i-th node in the graph, and d is the feature dimension;

[0144] S45. Output the constructed directed heterogeneous graph structure G, adjacency matrix A and node feature matrix X V , for subsequent graph convolutional neural network model processing and calling.

[0145] By mapping abnormal data into a directed heterogeneous graph, this paper defines the graph structure between the vehicle identification, charging time, charging location, and amount fields, and constructs an adjacency matrix and a node feature matrix, providing a rigorous structural foundation for subsequent graph convolution operations. This graph modeling approach enables the system to leverage the structural relationships within the data for deep learning, mining contextual features and improving the expressiveness of subsequent correction models.

[0146] In this embodiment, the directed isomeric graph G=(V, E, Φ, Ψ) is obtained by defining a graph structure, where:

[0147] V is the set of all nodes in the graph, satisfying V = V v ∪V t ∪V l ∪V a ;

[0148] V v ={v i} is the vehicle identification node set;

[0149] V t ={t i} is the charging time node set;

[0150] V l ={l i} is the charging location node set;

[0151] V a ={a i} is the charging amount node set;

[0152] is a set of directed edges;

[0153] Φ:V→{Vehicle,Time,Location,Amount} is the node type mapping function;

[0154] Ψ:E→{(Vehicle→Time),(Time→Location),(Location→Amount)} is the edge type mapping function.

[0155] This paper further refines the definitions of node types and edge relationships in a directed heterogeneous graph, specifying the directed connection logic between four types of nodes: vehicle identification, time, location, and amount, to construct a stable and universal graph structure template. This design ensures the logical integrity of node information transmission during graph neural network training, making embedded features more semantically expressive and enhancing the model's fit and transferability under complex data relationships.

[0156] In this embodiment, the directed edge relationship is through the abnormal record x i =[v i ,t i ,l i ,a i ] to obtain the graph structure including:

[0157] From the vehicle identification node v i Point to charging time node t i , define the directed edge

[0158] From the charging time node t i Point to charging location node l i , define the directed edge

[0159] From charging location node l i Point to the charging amount node a i , define the directed edge

[0160] This paper proposes a node embedding method based on a graph convolutional neural network. This method combines adjacency matrices with node feature matrices for layer-by-layer propagation. It learns the contextual dependencies of nodes in the graph structure through a fusion mechanism of multi-relationship weights and edge weights. Based on the embedded vectors, a multi-layer perceptron outputs predicted correction values. This method can respond specifically to field-level correction requests for abnormal data, enabling deep semantic generation of correction suggestions and significantly improving the accuracy of anomaly correction.

[0161] In this embodiment, the S5 specifically includes:

[0162] S51. On the directed heterogeneous graph G = (V, E, Φ, Ψ), combine the adjacency matrix A and the node feature matrix X V , define the l-th layer embedding output of the graph convolutional neural network as Initialize embedding to H (0) =X V ,in is the node representation matrix of the 0th layer of the graph convolutional neural network, d l is the dimension of the node embedding in the lth layer, |V| is the total number of nodes;

[0163] S52. Perform graph convolution operation and define the propagation function of the l+1th layer of the graph convolutional neural network based on the embedding result of the lth layer:

[0164]

[0165] in, represents the embedding representation of the l-th layer node, is the weight matrix of the lth layer, is the bias term, I is the unit matrix, which means the introduction of self-connection edges. for The degree matrix of σ(·) is the activation function, which takes ReLU, R is the relationship set in the heterogeneous graph, including edge types, is the adjacency matrix of relationship type r, A r The degree matrix of is the weight matrix of relationship type r in layer l, Fusion weights for heterogeneous relations;

[0166] S53, repeat the graph convolution propagation function to perform L-layer embedding propagation, and finally output the graph embedding matrix:

[0167]

[0168] where d L is the final graph embedding dimension, Z is the graph embedding matrix of all nodes;

[0169] S54. For each abnormal data record x i =[v i ,t i ,l i ,a i ], extract x from the graph embedding matrix Z i The corresponding four node embedding vectors Splice them into a structured graph representation and input them into the multi-layer perceptron prediction model to obtain the prediction correction value vector:

[0170]

[0171] Among them, ‖ represents the vector concatenation operation, Represents record x i The structured graph representation of For record x i The forecast revision of These are the predicted correction results of the vehicle identification, charging time, charging location and charging amount fields;

[0172] S55. Output the predicted correction value set of all abnormal records for subsequent field error judgment and abnormality correction.

[0173] This paper proposes a field-level error calculation and credibility determination mechanism. This uses a structured graph to represent prediction results and original anomaly data for field discrepancy analysis. It then sets specific error thresholds and correctability conditions to accurately assess the effectiveness and safety of corrections. By applying differentiated judgment logic to different fields, this mechanism effectively prevents incorrect or overcorrection errors, improving the reliability of correction results and the overall fault tolerance of the system.

[0174] In this embodiment, S6 specifically includes:

[0175] S61. For each abnormal record x i =[v i ,t i ,l i ,a i ] and its forecast revision value Calculate the absolute error between each field and construct a field-level error vector:

[0176]

[0177] in, They are defined as follows:

[0178] The value is 1 when the vehicle identifications are different, otherwise it is 0;

[0179] is the absolute error of the charging time field;

[0180] The value is 1 when the charging location is different, otherwise it is 0;

[0181] is the absolute error of the charge amount field;

[0182] is the indicator function;

[0183] S62, set the credibility judgment rule vector θ = [θ v ,θ t ,θ l ,θ a ], where ∈{0,1} is the vehicle identification correctability condition, is the maximum allowable error of charging time, θ l ∈{0,1} is the correctability condition of the charging location, The maximum permissible error in the charge amount;

[0184] According to the field-level error vector δ i With the credibility judgment rule θ, we judge whether the abnormal data of item i meets the correction condition and define the field credibility judgment function η(x i ):

[0185]

[0186] S63, for satisfying η(x i )=1, and construct the corrected data record x′ i =[v′ i ,t′ i ,l′ i ,a′ i ],in:

[0187]

[0188] S64. All records that meet the correction conditions and complete field replacement are combined into a corrected data set, and the corrected data set is output for the network charging data settlement center to update the original abnormal record.

[0189] This invention enables the automatic transmission and updating of corrected data, directly replacing the original abnormal record with the corrected data, ensuring data consistency and real-time performance within the settlement system. This function closes the data processing loop, reduces manual review and reprocessing steps, and improves the operational efficiency and automation level of the networked toll collection system, demonstrating its practical value and engineering scalability.

[0190] Example 1:

[0191] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the networked toll collection system under the jurisdiction of a certain highway operation and management unit. The coverage area of ​​the operation unit includes multiple highway sections between Nanjing and Changzhou in Jiangsu Province, with a total of 35 operating toll stations, an average daily vehicle traffic of about 210,000 vehicles, and an average monthly billing record of more than 6 million. The data collected by the system mainly include: vehicle identification (license plate number), charging time (accurate to seconds), charging location (entry / exit code), and charging amount (yuan). In the actual operation process, the system often has the following abnormal problems: some vehicles have inconsistent travel paths and time logic, such as Nanjing exit but Changzhou entry; some records have zero or abnormally low amounts, and their travel paths cannot be explained; some records have problems such as garbled license plates and missing timestamps. These abnormal data seriously interfere with subsequent clearing and settlement and financial verification. Long-term reliance on manual verification consumes a lot of manpower, and the accuracy is not high, which cannot meet the rapidly growing data processing needs.

[0192] During the implementation of the present invention, all networked toll collection data collected for this road section between October 1, 2024, and October 31, 2024, totaling approximately 6.528 million records, were first processed. Preprocessing operations such as data cleaning, normalization, and missing value filling were first performed on this dataset, and fields such as charging time, amount, and location code were standardized. Subsequently, the multi-granularity weighted isolation forest model proposed in the present invention was used to calculate anomaly scores for all data. Feature subspaces of three different granularities were constructed: vehicle identification, time + location combination, and amount + time combination. Weights were set for each granularity subspace to be 0.5, 0.3, and 0.2, respectively.

[0193] The isolation forest model generates the average path length for each record in each granularity subspace and calculates a final anomaly score using a weighted scoring formula. Based on this, a dynamic quantile threshold strategy is employed, setting the anomaly score above the 95th percentile as the anomaly criterion. Ultimately, approximately 127,000 anomalous data items, representing 1.94% of the total, were identified. Typical issues covered by these anomalies include: time inversion, zero amount, location jumps, and license plate recognition failures. Among these, 14,000 records included "zero amount but distance exceeding 100 kilometers," nearly 32,000 included "departure time earlier than arrival time," and over 98,000 included "missing or garbled license plate fields."

[0194] After identifying abnormal data, the proposed graph convolutional neural network method is used to correct the abnormal data. When constructing the graph structure, the vehicle identification, time, location, and amount fields are constructed as four types of nodes. The graph structure is established through the co-occurrence relationships between the fields, and graph embedding training is performed using a graph neural network. Each abnormal record is broken down into four nodes and connected to form a path. The embedding model is propagated through a three-layer GNN to obtain a context vector representation for each node. The four vectors are concatenated and fed into a multi-layer perceptron correction network to predict the corrected field for the vehicle. The final prediction results undergo field-level error calculation, are filtered using credibility judgment rules, and then a field replacement operation is performed. Taking the "garbled license plate" type anomaly as an example, the system's license plate recovery accuracy using the model correction is 96.2%. For records with an "amount of 0" and a path exceeding 100 kilometers, the model's corrected amount error is within ±3 yuan, accounting for over 88% of the records.

[0195] The entire process was automatically completed by the model, with a total processing time of 38 minutes. Manually verifying the same number of records would typically require 3-5 people working for more than two days. Ultimately, 112,000 valid records, representing 88.2% of the total anomalies, were corrected and returned to the settlement system using this method. The remaining anomalies were flagged for manual review due to insufficient credibility. The following table shows typical anomaly types, number of anomalies, accuracy after correction, and system processing time:

[0196] Table 1 Statistics of abnormality detection and correction results of a certain highway network toll collection system in Jiangsu in October 2024

[0197]

[0198]

[0199] This invention significantly reduces the need for manual intervention, improves the accuracy of anomaly identification and the real-time nature of corrections, and demonstrates strong feasibility and engineering value in the networked highway toll collection system. Operators have reported that the corrected data has reduced clearing and reconciliation disputes by over 98%, significantly improving customer satisfaction and operational efficiency.

[0200] It can be seen from the above embodiments that the present invention not only effectively solves the problem that the existing technology is difficult to automatically process complex abnormal data in the networked charging data scenario, but also significantly improves the overall processing efficiency and correction accuracy, and has good practical application prospects and promotion value.

[0201] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for detecting and correcting anomalies in online charging data based on a machine learning model, characterized in that: It includes the following steps: S1. Collect network toll data and perform preprocessing; S2. Based on the preprocessed toll data, construct and train a multi-granularity weighted isolation forest model, randomly obtain multiple isolation trees, count the average path length of each toll data in each isolation tree, and calculate the corresponding anomaly score; S3. Adopt a dynamic quantile strategy to determine the anomaly threshold. If the anomaly score is greater than the anomaly threshold, determine the toll data as abnormal data and output the abnormal data set; S4. Map the abnormal data set to graph-structured data, use vehicle identification, toll time, toll location, and toll amount as node fields, establish directed edges between nodes through the co-occurrence relationship between fields, and form a directed heterogeneous graph; S5. Based on the directed heterogeneous graph, construct a graph convolutional neural network model, perform graph embedding training on the nodes through the node aggregation and edge feature propagation mechanisms, obtain the structured graph representation of each abnormal record, and output the corresponding prediction correction value based on the graph representation; S6. Perform field-level error calculation on the prediction correction value and the abnormal data, and combine the credibility judgment rule to screen the abnormal records that meet the correction conditions, and perform the field replacement operation to generate the corrected data set; S7. Send the corrected data set to the network toll data settlement center, replace the original abnormal records, and complete the real-time update of the toll records.

2. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 1, characterized in that: The toll data includes vehicle identification, toll time, toll location, and toll amount, and the preprocessing includes data cleaning, normalization processing, and missing value filling.

3. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 1, characterized in that: The specific content of S2 includes: S21, extract feature fields from the pre-processed charging data, and construct a feature vector set X = {x1, x2, ..., x n }, where x i =[v i ,t i ,l i ,a i ], is the feature vector corresponding to the i-th charging record, n is the total number of charging records, v i is the vector corresponding to the vehicle identification field in the i-th record, t i is the normalized value of the charging time field in the i-th record, l i is the code value of the charging location field in the i-th record, a i is the normalized value of the charge amount field in the i-th record; S22, based on the feature vector set X, construct multiple feature granularity subspaces, and define the granularity set as G = {g1, g2, ..., g k }, where k is the number of granular subspaces, is the jth granularity subspace composed of a field combination, wherein the field combination consists of a vehicle identification field, a charging time field, a charging location field, and a charging amount field; S23, in each granularity subspace g j In the random subspace sampling method, m j isolated trees, forming an isolated tree set Where T jp is the granularity subspace g j The pth isolated tree constructed under j is the granularity subspace g j the number of isolated trees; S24. For each eigenvector x i In each isolated tree T jp Get the path length, which is recorded as h(x i ,T jp ), calculate the particle size g j The average path length under , and weight the average path length of each particle size: in, is the particle size g j The average path length of the next i-th record, w j is the weight coefficient of the j-th granularity subspace, is the weighted average path length of the i-th record at all granularities, satisfying w j ∈[0,1], and S25, according to the weighted average path length Calculate the anomaly score of the i-th record: Among them, s(x i ) is the anomaly score of the ith record, c(n) is the normalization constant that depends on the total number of charging records n, H(n) is the nth harmonic number, H(n-1) is the n-1th harmonic number, and γ is the Euler constant, which is 0.5772.

4. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 3 is characterized in that: The specific content of S23 includes: S231, for the granularity subspace g j The feature set in , randomly extract samples from the feature vector set X without replacement to form a training sample set The number of training samples is Where n is the total number of charging records; S232. In each isolated tree T jp During the construction process, the following steps are recursively performed until the maximum depth d is reached. max Or the number of leaf node samples is 1: S2321. Randomly select a dimension f∈g from the feature dimension of the current node j As a dividing dimension; S2322. Randomly select a split point q within the value range of the feature dimension f, and divide the sample into a left subtree and a right subtree according to f(x) < q; S2323. Continue to perform the splitting operation on the left and right subtrees to construct a tree structure; S233, set maximum depth Ensure that the depth of the tree structure is adapted to the sample size to improve isolation efficiency; S234, all the isolated trees T that have been built jp The isolation forest set F stored at the corresponding granularity level j It is used for subsequent path length calculation.

5. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 1, characterized in that: The specific content of S3 includes: S31. Get all eigenvectors x i The set of abnormal scores is denoted as S = {s(x1),s(x2),…,s(x n )}, where s(x i ) is the anomaly score of the i-th record, and n is the total number of charging records; S32. Sort the abnormal score set S in non-descending order according to the numerical value to obtain the sorted score sequence S ′ ={s ′ 1,s ′ 2,…,s ′ n }, satisfying s ′ 1≤s ′ 2≤…≤s ′ n ; S33. Set the target quantile level q ∈ (0, 1), use the quantile level q as the percentage index of the corresponding position of the threshold, and calculate its corresponding quantile position index P; Among them, q is the target quantile level, N is the number of scoring samples, s′ i is the i-th score after sorting, log(N+1) is the regularization factor, which is used to balance the impact of sample size on the quantile position. is the sum of abnormal scores, max(S′) is the maximum value in the score sequence, min(S′) is the minimum value in the score sequence, is the floor function; S34. According to the position index P, select the Pth element from the sorted scoring sequence as the abnormality judgment threshold, set to T q , each eigenvector x i The anomaly score s(x i ) is compared with the threshold, if s(x i )>T q , then the data is judged to be abnormal data, and the abnormal data set X that meets the conditions is output abn .

6. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 1, characterized in that: The specific content of S4 includes: S41. Extract abnormal data set X abn ={x i ∣s(x i )>T q }, and define a directed heterogeneous graph G, where x i =[v i ,t i ,l i ,a i ], is the feature vector corresponding to the i-th charging record, n is the total number of charging records, v i is the vector corresponding to the vehicle identification field in the i-th record, t i is the normalized value of the charging time field in the i-th record, l i is the code value of the charging location field in the i-th record, a i is the normalized value of the charge amount field in the i-th record, s(x i ) is the abnormal score of the i-th record, T q is the threshold; S42, according to each abnormal record x i =[v i ,t i ,l i ,a i ], establish directed edge relationships in the graph and group all directed edges into edge sets in For directed edges, further generate the final directed heterogeneous graph G; S43. Constructing the adjacency matrix between nodes The definition is as follows: Among them, A u,v is the connection relationship from node u to node v, |V| is the total number of nodes; S44. Construct node feature matrix Each of these lines is the feature vector of the i-th node in the graph, and d is the feature dimension; S45. Output the constructed directed heterogeneous graph structure G, adjacency matrix A and node feature matrix X V , for subsequent graph convolutional neural network model processing and calling.

7. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 6, characterized in that: The directed heterogeneous graph G = (V, E, Φ, Ψ) is obtained through the definition of the graph structure, where: V is the set of all nodes in the graph, satisfying V = V v ∪V t ∪V l ∪V a ; V v ={v i } is the vehicle identification node set; V t ={t i } is the charging time node set; V l ={l i } is the charging location node set; V a ={a i } is the charging amount node set; is a set of directed edges; Φ: V → {Vehicle, Time, Location, Amount} is the node type mapping function; Ψ: E → {(Vehicle → Time), (Time → Location), (Location → Amount)} is the edge type mapping function.

8. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 6, characterized in that: The directed edge relationship is through the exception record x i =[v i ,t i ,l i ,a i ] to obtain the graph structure including: From the vehicle identification node v i Point to charging time node t i , define the directed edge From charging time node t i Point to charging location node l i , define the directed edge From charging location node l i Point to the charging amount node a i , define the directed edge 9. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 1, characterized in that: The specific content of S5 includes: S51. On the directed heterogeneous graph G = (V, E, Φ, Ψ), combine the adjacency matrix A and the node feature matrix X V , define the l-th layer embedding output of the graph convolutional neural network as Initialize embedding to H (0) =X V ,in is the node representation matrix of the 0th layer of the graph convolutional neural network, d l is the dimension of the node embedding in the lth layer, |V| is the total number of nodes; S52. Perform graph convolution operation, and based on the embedding result of the l-th layer, define the propagation function of the (l + 1)-th layer of the graph convolutional neural network: in, represents the embedding representation of the l-th layer node, is the weight matrix of the lth layer, is the bias term, I is the unit matrix, which means the introduction of self-connection edges. for The degree matrix of σ(·) is the activation function, which takes ReLU, R is the relationship set in the heterogeneous graph, including edge types, is the adjacency matrix of relationship type r, A r The degree matrix of is the weight matrix of relationship type r in layer l, Fusion weights for heterogeneous relations; S53. Repeat the graph convolution propagation function for L-layer embedding propagation, and finally output the graph embedding matrix: where d L is the final graph embedding dimension, Z is the graph embedding matrix of all nodes; S54. For each abnormal data record x i =[v i ,t i ,l i ,a i ], extract x from the graph embedding matrix Z i The corresponding four node embedding vectors Splice them into a structured graph representation and input them into the multi-layer perceptron prediction model to obtain the prediction correction value vector: Among them, ‖ represents the vector concatenation operation, Represents record x i The structured graph representation of For record x i The forecast revision of These are the predicted correction results of the vehicle identification, charging time, charging location and charging amount fields; S55. Output the prediction correction value set of all abnormal records for subsequent field error judgment and anomaly correction calls.

10. The method for detecting and correcting anomalies in online charging data based on a machine learning model according to claim 1, characterized in that: The specific content of S6 includes: S61. For each abnormal record x i =[v i ,t i ,l i ,a i ] and its forecast revision value Calculate the absolute error between each field and construct a field-level error vector: in, They are defined as follows: The value is 1 when the vehicle identifications are different, otherwise it is 0; is the absolute error of the charging time field; The value is 1 when the charging location is different, otherwise it is 0; is the absolute error of the charge amount field; is the indicator function; S62, set the credibility judgment rule vector θ = [θ v ,θ t ,θ l ,θ a ], where ∈{0,1} is the vehicle identification correctability condition, is the maximum allowable error of charging time, θ l ∈{0,1} is the correctability condition of the charging location, The maximum permissible error in the charge amount; According to the field-level error vector δ i With the credibility judgment rule θ, we judge whether the abnormal data of item i meets the correction condition and define the field credibility judgment function η(x i ): S63, for satisfying η(x i )=1, and construct the corrected data record x′ i =[v′ i ,t′ i ,l′ i ,a′ i ],in: S64. Combine all records that meet the correction conditions and complete field replacement into the corrected data set, and output the corrected data set for the network toll data settlement center to update the original abnormal records.