Multi-path toll splitting method based on big data processing and related device
Through big data processing technology and multi-path toll splitting method, the problems of large errors and low accuracy in traditional toll calculation methods are solved, and the precise splitting of expressway tolls and high-precision prediction of future income are achieved.
Patent Information
- Application Number
- CN202510219517.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional highway toll calculation method has problems such as large allocation errors, low accuracy, and difficulty in predicting future income.
The multi-path toll splitting method based on big data processing is adopted to accurately split and full life cycle revenue forecast by identifying and matching vehicle towing paths, combining Markov path selection model and gradient descent hybrid algorithm.
It realizes accurate splitting of tolls and high-precision prediction of charging income, reduces errors and improves the accuracy of investment benefit calculation.
Smart Images

Figure CN120048012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of highway toll calculation, and in particular to a multi-path toll splitting method and device based on big data processing, and a computing device. Background Art
[0002] In the field of highway toll calculation, the traditional method is to calculate the toll revenue according to the toll standard by vehicle type, and then clear it through special software. It puts forward high requirements for the automatic classification of vehicles and the identification of driving trajectories, the continuous operation of the toll monitoring system, and the comprehensive quality of the calculation personnel. It will also produce certain allocation errors and a large coordination workload. Moreover, it can only predict the past but not the future income. It is difficult to accurately calculate the toll revenue of the past and future years, and the error is large and the accuracy is low. Summary of the invention
[0003] To solve the above problems, the present invention discloses a multi-path toll splitting method and device based on big data processing, and a computing device, so as to accurately split multi-path tolls and predict toll revenue for the entire life cycle through big data processing methods, with lower error and higher accuracy, exploring new ways for project investment benefit calculation.
[0004] According to one aspect of the present invention, a multi-path toll splitting method based on big data processing is provided, comprising:
[0005] Step S1, identifying and matching vehicle travel paths according to license plate information, highway network topology information, and toll station geographic location information to obtain vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set, and a disputed unmatched path set;
[0006] Step S2, according to the actual mileage of the vehicle, the differentiated charging standards of each road section and the vehicle path information, the tolls of the paths in the undisputed path set and the disputed matched path set are split to obtain a first split fee; wherein, for vehicles passing through multiple paths in the road network, based on the historical traffic distribution and the congestion information of the adjacent paths, the passing probability of the vehicle on the vehicle path information is predicted according to the Markov path selection model, and according to the passing probability, the road section service level index and the difference in the distance of the alternative paths, the first split fee is adjusted to obtain the second split fee data;
[0007] Step S3, performing target optimization on the second split cost data according to a gradient descent hybrid algorithm to obtain third split cost data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data;
[0008] Step S4, calculating an error index based on the third split cost data and the actual charging data, and predicting the charging income and total income for each year of the whole life cycle based on the error index and the historical charging data.
[0009] In an optional manner, the license plate information includes the license plate number, vehicle type, vehicle axle number and vehicle exhaust emission standard;
[0010] The highway network topology information includes the section number, section length, section connection relationship, section type, section design speed, section slope and section curvature;
[0011] The toll station geographic location information includes the toll station number, the toll station geographic coordinates, the toll station service range, the number of toll station lanes and the toll station ETC / MTC ratio.
[0012] In an optional manner, adjusting the first split cost to obtain second split cost data further includes:
[0013] Step S21, training is performed based on the historical traffic data of the vehicle, and the state transfer matrix and emission matrix of the HMM are constructed by using the maximum likelihood estimation method; wherein the state space is a set of road sections; the state transfer probability is determined based on the historical traffic data and the connectivity of the road sections, indicating the probability of a vehicle transferring from one road section to another; the observation space is the multimodal perception data of the vehicle, which is used to infer the current state of the vehicle; the emission probability is calculated based on the matching degree between the vehicle position and the road section;
[0014] Step S22, dynamically adjusting the state transition probability of the HMM according to the congestion level of the adjacent paths;
[0015] Step S23, decoding the HMM according to the Viterbi algorithm to predict the passing probability of the vehicle on each possible path;
[0016] Step S24, adjusting the first split cost by a correction factor according to the road section service level index and the difference in the alternative path distance to obtain second split cost data.
[0017] In an optional manner, performing target optimization on the second split cost data according to the gradient descent hybrid algorithm to obtain the third split cost data further includes:
[0018] Step S31, defining an objective function for measuring the error between the split fee and the actual charging data;
[0019] Step S32, calculating the gradient of the objective function with respect to the splitting cost, wherein the gradient represents the change direction and magnitude of the objective function under the current splitting cost;
[0020] Step S33, updating the first-order moment estimate and the second-order moment estimate according to the Adam algorithm; wherein the first-order moment estimate is the average exponential shift of the gradient, and the second-order moment estimate is the average exponential shift of the gradient square;
[0021] Step S34, updating the splitting cost data according to the update rule of the Adam algorithm until the objective function converges or reaches a preset number of iterations, and outputting the optimized third splitting cost data.
[0022] In an optional manner, the predicting of charging income and total income for each year of the life cycle according to the error index and historical charging data further includes:
[0023] Step S51, model the historical toll data according to the seasonal SARIMA model, and solve the model order (p, d, q) or (p, d, q)(P, D, Q) s , so that the AIC or BIC value of the seasonal SARIMA model is minimized;
[0024] Step S52: predicting the total income of the entire life cycle according to the seasonal SARIMA model.
[0025] In an optional manner, the calculation formula for the alternative path distance difference is:
[0026]
[0027] Among them, d 1i and d 2i are the lengths of the i-th type of road on the current path and the alternative path respectively; w i is the weight of the i-th type of road.
[0028] In an optional manner, the solution model order is (p, d, q) or (p, d, q)(P, D, Q) s , minimizing the AIC or BIC value of the seasonal SARIMA model further includes:
[0029] Repeat the first-order difference of the historical toll data, record the order d of the difference, and repeat the seasonal difference of the historical toll data, record the order D of the seasonal difference;
[0030] According to the autocorrelation function and the partial autocorrelation function, the AR and MA orders (p, q, P, Q) are preliminarily estimated, and multiple SARIMA models are constructed according to the orders (p, q, P, Q);
[0031] Fit various SARIMA models according to the historical charging data, calculate the AIC value and BIC value of each SARIMA model, and select the SARIMA model with the smallest AIC or BIC value as the best model.
[0032] In an optional manner, the calculation formula of the AIC value is:
[0033]
[0034] Among them, Likelihood(L) is the likelihood of historical toll data; K is the number of parameters of the model; ∫α(x)*f(x)dx is the fitting error of toll data; ∑[γ j *g(t j ) represents the contribution of a preset road section or toll booth; k The weight of the preset road section or toll booth; represents the probability contribution of a preset road section or toll station; k is the traffic volume of the preset road section or toll station; h(z k ) is z k The service level of k ) is z k The probability density function of γ j is the weight of the preset road section or toll station j; g(t j ) is t j The degree of traffic congestion; j is the characteristic variable of road section or toll station j; α(x) is the weight function of toll data x; f(x) is the normal distribution function of toll data x;
[0035] The calculation formula of the BIC value is:
[0036] BIC=-2×log(∫α(x)f(x)dx)+log(h(z k ))
[0037] Where log(∫α(x)f(x)dx) is the fitting error of the charging data.
[0038] According to another aspect of the present invention, a multi-path toll splitting device based on big data processing is provided, comprising:
[0039] A path identification module is used to identify and match the vehicle travel path according to the license plate information, the highway network topology information and the toll station geographical location information to obtain the vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set and a disputed unmatched path set;
[0040] A path splitting module is used to split the tolls of the paths in the undisputed path set and the disputed matched path set according to the actual mileage of the vehicle, the differentiated charging standards of each road section and the vehicle path information to obtain a first split fee; wherein, for vehicles passing through multiple paths in the road network, based on the historical traffic distribution and the congestion information of the adjacent paths, the passing probability of the vehicle on the vehicle path information is predicted according to the Markov path selection model, and the first split fee is adjusted according to the passing probability, the road section service level index and the difference in the distance of the alternative paths to obtain the second split fee data;
[0041] A target optimization module, used for performing target optimization on the second split cost data according to a gradient descent hybrid algorithm to obtain third split cost data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data;
[0042] The error evaluation module is used to calculate the error index based on the third split cost data and the actual charging data, and predict the charging income and total income of each year of the whole life cycle based on the error index and historical charging data.
[0043] According to another aspect of the present invention, there is provided a computing device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus;
[0044] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned multi-path toll splitting method based on big data processing.
[0045] According to the scheme provided by the present invention, it includes: step S1, identifying and matching the vehicle travel path according to the license plate information, the highway network topology information and the geographical location information of the toll station to obtain the vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set and a disputed unmatched path set; step S2, according to the actual mileage of the vehicle, the differentiated charging standards of each section and the vehicle path information, the tolls in the undisputed path set and the disputed matched path set are split to obtain the first split fee; wherein, for vehicles traveling on multiple paths in the road network, based on the historical flow distribution and the congestion of adjacent paths, Information, predict the passing probability of the vehicle on the vehicle path information according to the Markov path selection model, adjust the first split fee to obtain the second split fee data according to the passing probability, the road section service level index and the difference in the distance of the alternative path; step S3, target optimization of the second split fee data according to the gradient descent hybrid algorithm to obtain the third split fee data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data; step S4, calculate the error index according to the third split fee data and the actual charging data, and predict the toll revenue and total revenue of each year of the whole life cycle according to the error index and the historical charging data. The present invention uses a big data processing method to accurately split multi-path tolls and predict the toll revenue of the whole life cycle, with low error and high accuracy, exploring new ways for project investment benefit measurement.
[0046] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0048] Figure 1 A schematic diagram showing a flow chart of a multi-path toll splitting method based on big data processing according to an embodiment of the present invention;
[0049] Figure 2 A schematic diagram of a road network structure according to an embodiment of the present invention is shown;
[0050] Figure 3 A schematic diagram of the framework of a multi-path toll splitting device based on big data processing according to an embodiment of the present invention is shown;
[0051] Figure 4 A schematic diagram of the structure of a computing device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0052] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.
[0053] Figure 1 The flowchart of the multi-path toll splitting method based on big data processing according to an embodiment of the present invention is shown. Figure 1 As shown, the following steps are included:
[0054] Step S1, identifying and matching vehicle travel paths according to license plate information, highway network topology information and toll station geographic location information to obtain vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set and a disputed unmatched path set.
[0055] In this embodiment, by combining multi-source heterogeneous data (license plates, road networks, toll stations), compared with a single data source, it is not only possible to distinguish vehicles entering and exiting the same toll station, but it is also easier to determine whether the vehicle has traveled a complete path. Especially in the case of complex road networks and multiple optional paths, it is more effective to identify the possible set of vehicle travel paths. By distinguishing between "undisputed paths", "disputed matched paths" and "disputed unmatched paths", it is convenient to adopt different processing strategies for different situations, thereby improving the rationality of toll splitting.
[0056] Specifically, a road network structure (topology map) including node (toll station, road section connection point) and edge (road section) information is constructed, and the road section information is calibrated to ensure the accuracy of the road section length and connection relationship. The toll station information (number, coordinates, service range, etc.) is associated with the road network structure. Use algorithms such as A* combined with road network topology information to calculate the shortest path between the vehicle's starting point (entrance toll station) and the end point (exit toll station). Analyze the historical vehicle traffic records and count the traffic frequencies of different paths as the prior probability of path selection. According to the restrictions of the actual road (such as one-way streets, prohibited sections), etc., unreasonable paths are excluded. Since the vehicle may choose a non-shortest path in actual situations, it is necessary to identify multiple possible paths or select the top N shortest paths. In this embodiment, the undisputed path set refers to the only reasonable path identified by the above algorithm. For example, when a vehicle is driving on a highway, there is only one road between the entrance and exit toll stations. The disputed matched path set refers to identifying multiple possible paths, but determining the path that the vehicle is most likely to choose based on historical data. The disputed unmatched path set refers to the identification of multiple possible paths, and it is impossible to determine the path actually selected by the vehicle.
[0057] In an optional manner, the license plate information includes the license plate number, vehicle type, vehicle axle number and vehicle exhaust emission standard;
[0058] The highway network topology information includes the section number, section length, section connection relationship, section type, section design speed, section slope and section curvature;
[0059] The toll station geographic location information includes the toll station number, the toll station geographic coordinates, the toll station service range, the number of toll station lanes and the toll station ETC / MTC ratio.
[0060] In this embodiment, the vehicle type and the number of axles affect the charging standard of the vehicle, which can accurately determine the rate that should be applied to the vehicle and avoid incorrect charges. Vehicle exhaust emission standards are used to implement differentiated charging policies, such as charging higher tolls for high-emission vehicles to encourage the use of environmentally friendly vehicles. The license plate number is the unique identifier of the vehicle and is used to verify the legality of the vehicle and prevent toll evasion. The section length information can improve the accuracy of toll calculations, and the section connection relationship can help identify all possible travel paths. The section types (such as expressways and first-class highways) correspond to different charging standards. The design speed, slope and curvature of the section affect the vehicle's driving speed and comfort, thereby affecting the driver's path selection, and can more accurately predict the vehicle's travel probability. The geographic coordinates of the toll station are used to determine the vehicle's entry and exit toll stations. The number of toll station lanes and the ETC / MTC ratio reflect the toll station's capacity and are used to evaluate the congestion of the toll station and predict the waiting time of the vehicle, thereby affecting the path selection probability. The toll station service range identifies the possible travel path of the vehicle. Figure 2 As shown in the figure, each node (N1, N2, etc.) represents a toll station, including information such as number, geographical coordinates, and service range, such as N1 node: number 1, coordinates (0,0), service range: 10km. Each edge (such as E1, E2) represents a road section, including information such as road section number and length (such as number 5, length 7km). Among them, the solid line represents the set of undisputed paths, the dotted line represents the set of disputed matched paths, and the dot-dash line represents the set of disputed unmatched paths.
[0061] Step S2, according to the actual mileage of the vehicle, the differentiated charging standards of each road section and the vehicle path information, the tolls of the paths in the undisputed path set and the disputed matched path set are split to obtain a first split fee; wherein, for vehicles passing through multiple paths in the road network, based on the historical traffic distribution and congestion information of adjacent paths, the passing probability of the vehicle on the vehicle path information is predicted according to the Markov path selection model, and according to the passing probability, the road section service level indicator and the difference in the distance of the alternative paths, the first split fee is adjusted to obtain the second split fee data.
[0062] In this embodiment, the tolls of vehicles on each road section can be calculated more accurately by using the actual mileage of the vehicle and the differentiated charging standards of each road section. For vehicles traveling on multiple paths in the road network, the tolls can be allocated more reasonably by predicting the vehicle's travel probability through the Markov path selection model. Based on the historical traffic distribution and congestion information of adjacent paths, the allocation of tolls can be dynamically adjusted to improve the rationality of the splitting results. Through the road section service level indicators and the distance differences of alternative paths, the tolls of vehicles can be more comprehensively evaluated to improve the fairness of the splitting.
[0063] Specifically, for the undisputed path set, the toll of each road section is calculated according to the actual mileage of the vehicle and the charging standards of each road section to obtain the first split fee. For the disputed matched path set, the toll of each road section is calculated according to the actual mileage of the vehicle and the charging standards of each road section to obtain the first split fee.
[0064] The road segments are regarded as state space and each road segment is regarded as a state. According to the historical traffic distribution and the connectivity of the road segments, the probability of a vehicle transferring from one road segment to another is calculated.
[0065] The vehicle's multimodal perception data (such as license plate recognition) is used as the observation space to infer the vehicle's current state. The emission probability is calculated based on the matching degree between the vehicle's position and the road section.
[0066] The traffic monitoring system is used to obtain real-time congestion information of adjacent paths, and the state transition probability of the Markov model is dynamically adjusted according to the congestion information. The Viterbi algorithm is used to decode the Markov model and predict the probability of vehicles passing through each possible path.
[0067] The toll is adjusted according to the service level indicators of the road section (such as the design speed, slope, curvature, etc. of the road section), the distance difference between the current path and the alternative path is calculated, and the toll is adjusted by the correction factor to obtain the second split fee data.
[0068] In an optional manner, adjusting the first split cost to obtain second split cost data further includes:
[0069] Step S21, training is performed based on the historical traffic data of the vehicle, and the state transfer matrix and emission matrix of the HMM are constructed by using the maximum likelihood estimation method; wherein the state space is a set of road sections; the state transfer probability is determined based on the historical traffic data and the connectivity of the road sections, indicating the probability of a vehicle transferring from one road section to another; the observation space is the multimodal perception data of the vehicle, which is used to infer the current state of the vehicle; the emission probability is calculated based on the matching degree between the vehicle position and the road section;
[0070] Step S22, dynamically adjusting the state transition probability of the HMM according to the congestion level of the adjacent paths;
[0071] Step S23, decoding the HMM according to the Viterbi algorithm to predict the passing probability of the vehicle on each possible path;
[0072] Step S24, adjusting the first split cost by a correction factor according to the road section service level index and the difference in the alternative path distance to obtain second split cost data.
[0073] For example, a highway consists of three sections A, B, and C. A vehicle entering from section A may pass through section B or section C. Therefore, it is necessary to adjust the fee split based on the vehicle's historical traffic data and real-time congestion. Collect the vehicle's historical traffic data on sections A, B, and C.
[0074] Construct an HMM model, where the state space is {A, B, C}, and the state transition probability is determined based on historical traffic data and road section connectivity, for example: P(A→B) = 0.6, P(A→C) = 0.4. If road section B is currently congested, then reduce P(A→B) (e.g., adjust to 0.4) and increase P(A→C) (e.g., adjust to 0.6) accordingly.
[0075] The Viterbi algorithm is used to decode the HMM model and predict the probability of a vehicle passing through section B and section C. The prediction results are: P(B) = 0.5, P(C) = 0.5.
[0076] The cost is adjusted by the correction factor according to the service level index of the road section and the difference in the distance of the alternative path. Assume that the service level of section B is higher, the service level of section C is lower, and the alternative path distance of section C is longer. Correction factor B = 1.2, correction factor C = 0.8.
[0077] The final cost breakdown is:
[0078] Cost B = 0.5 × 1.2 = 0.6
[0079] Cost C = 0.5 × 0.8 = 0.4.
[0080] In an optional manner, the calculation formula for the alternative path distance difference is:
[0081]
[0082] Among them, d 1i and d 2i are the lengths of the i-th type of road on the current path and the alternative path respectively; w i is the weight of the i-th type of road.
[0083] In this embodiment, the contribution of different types of roads to the actual travel distance is combined to further improve the accuracy of the distance difference between the current path and the alternative path. The distance difference of the alternative path can be used to split the cost more fairly, especially in the case of multiple optional paths.
[0084] For example, a truck can choose two routes from point A to point B:
[0085] Route 1 (current route): 100 km of expressways and 50 km of national roads;
[0086] Route 2 (alternative route): 80 km on expressways and 70 km on national roads;
[0087] Among them, the highway weight w 1 =1.0, national highway weight w 2 =0.5, dart weight w 3 =0.3.
[0088] Path 1: d 11 (Highway) = 100 km, d 12 (National Highway) = 50 km
[0089] Path 2: d 21 (Highway) = 80 km, d 22 (National Highway) = 70 km
[0090] Substituting into the above formula, we obtain the alternative path distance difference = 10 / 240≈0.042.
[0091] Assume that the first split cost is 100 yuan. According to the previous HMM model, the pass probabilities of path 1 and path 2 are 0.6 and 0.4 respectively. The original split costs are: path 1 = 60 yuan, path 2 = 40 yuan. Using the difference in the distance of the alternative paths as a correction factor, the costs can be adjusted as follows:
[0092] Corrected path 1 cost = 60 × (1-0.042) = 57.48 yuan
[0093] The corrected cost of path 2 = 40×(1+0.042×(60 / 40)) = 42.52 yuan.
[0094] In the above example, the cost of path 1 is slightly reduced and the cost of path 2 is slightly increased because the alternative path is shorter.
[0095] Step S3, performing target optimization on the second split cost data according to a gradient descent hybrid algorithm to obtain third split cost data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data.
[0096] In this embodiment, the gradient descent hybrid algorithm can effectively minimize the error (mean square error, mean absolute error, etc.) between the split result and the actual charging data, and adaptively learn the best fee splitting solution through iterative optimization without human intervention.
[0097] In an optional manner, performing target optimization on the second split cost data according to the gradient descent hybrid algorithm to obtain the third split cost data further includes:
[0098] Step S31, defining an objective function for measuring the error between the split fee and the actual charging data;
[0099] Step S32, calculating the gradient of the objective function with respect to the splitting cost, wherein the gradient represents the change direction and magnitude of the objective function under the current splitting cost;
[0100] Step S33, updating the first-order moment estimate and the second-order moment estimate according to the Adam algorithm; wherein the first-order moment estimate is the average exponential shift of the gradient, and the second-order moment estimate is the average exponential shift of the gradient square;
[0101] Step S34, updating the splitting cost data according to the update rule of the Adam algorithm until the objective function converges or reaches a preset number of iterations, and outputting the optimized third splitting cost data.
[0102] In this embodiment, the objective function may be in the form of mean square error (MSE), mean absolute error (MAE), etc., which is used to measure the error between the split fee and the actual charging data.
[0103] Calculate the gradient of the objective function with respect to the splitting cost. The gradient represents the direction and magnitude of change of the objective function under the current splitting cost.
[0104] Update the first-order moment estimate and the second-order moment estimate according to the Adam algorithm. The first-order moment estimate is the average exponential shift of the gradient: m t =β 1 ×m t-1 +(1-β 1 )×g t The second-order moment is estimated as the average exponential shift of the squared gradient: v t =β 2 ×v t-1 +(1-β 2 )×g t 2 Among them, β 1 and β 2 is the exponential decay rate, g t is the gradient at the current moment.
[0105] Update the split cost data according to the update rule of the Adam algorithm. First, perform bias correction on the first-order moment estimate and the second-order moment estimate:
[0106] m hat =m t / (1-β 1 t )
[0107] v hat =v t / (1-β 2t )
[0108] Then, update the split cost data:
[0109] fee = fee-learning rate *m hat / (sqrt(v hat )+ε)
[0110] Among them, learning rate is the learning rate, and ε is a small constant used to prevent the denominator from being zero.
[0111] Repeat until the objective function converges or reaches the preset number of iterations, and output the optimized third split cost data.
[0112] Step S4, calculating an error index based on the third split cost data and the actual charging data, and predicting the charging income and total income for each year of the whole life cycle based on the error index and the historical charging data.
[0113] In this embodiment, the prediction model includes a time series model (such as ARIMA), a regression model and a machine learning model (such as random forest, support vector machine, etc.) The prediction model is used to predict the charging income and total income of each year in the whole life cycle.
[0114] In an optional manner, the predicting of charging income and total income for each year of the life cycle according to the error index and historical charging data further includes:
[0115] Step S51, model the historical toll data according to the seasonal SARIMA model, and solve the model order (p, d, q) or (p, d, q)(P, D, Q) s , so that the AIC or BIC value of the seasonal SARIMA model is minimized;
[0116] Step S52: predicting the total income of the entire life cycle according to the seasonal SARIMA model.
[0117] In this embodiment, the seasonal SARIMA model is used to model the historical toll data, and the model order (p, d, q) and seasonal order (P, D, Q) are solved by minimizing the AIC or BIC value. s .
[0118] The fitting effect of the model is evaluated, the model with the smallest AIC or BIC value is selected, and the optimal seasonal SARIMA model is used to predict the toll revenue of the entire life cycle. Table 1 shows the toll revenue estimation table of a certain high-speed project during the operation period.
[0119] Table 1
[0120]
[0121] From the above table, we can see that the charging income of past and future years and the charging income of the entire life cycle can be predicted.
[0122] In this application, the new model predicts that the project's revenue in 2023 will be 52.74 million yuan, and the traditional method classification predicts that the revenue will be 53.72 million yuan, with a relative error of (53.72-52.74) / 53.72=1.82%. From the results, it can be seen that the overall error is only 1.82%, which is basically acceptable. The main reasons for the error are as follows:
[0123] First, various approximate values and weighted values produce certain errors;
[0124] 2. The International Standard Container Union's preferential vehicles (50% discount) are not included in the statistics;
[0125] 3. The toll errors caused by vehicles with unknown routes (charging based on the shortest route) are not counted;
[0126] 4. Free passage for emergency rescue vehicles (approved by the competent authorities) is not counted;
[0127] 5. The actual proportion of preferential passage for passenger cars and freight vehicles may not reach 5%;
[0128] 6. The ratio of free vehicles on the four major holidays may be slightly biased.
[0129] In an optional manner, the solution model order is (p, d, q) or (p, d, q)(P, D, Q) s , minimizing the AIC or BIC value of the seasonal SARIMA model further includes:
[0130] Repeat the first-order difference of the historical toll data, record the order d of the difference, and repeat the seasonal difference of the historical toll data, record the order D of the seasonal difference;
[0131] According to the autocorrelation function and the partial autocorrelation function, the AR and MA orders (p, q, P, Q) are preliminarily estimated, and multiple SARIMA models are constructed according to the orders (p, q, P, Q);
[0132] Fit various SARIMA models according to the historical charging data, calculate the AIC value and BIC value of each SARIMA model, and select the SARIMA model with the smallest AIC or BIC value as the best model.
[0133] In this embodiment, the model order can be estimated more accurately through the difference and autocorrelation function, and the number of models that need to be fitted can be reduced by preliminarily estimating the AR and MA orders, thereby reducing the calculation cost.
[0134] For example, the historical toll data of the highway is converted into a data series, and the first-order difference (d=1) and seasonal difference (cycle is 12 months) (D=1) are performed.
[0135] Draw ACF and PACF graphs, preliminarily estimate the AR and MA orders, and preliminarily estimate (p=1, q=1, P=1, Q=1) based on the characteristics of the ACF and PACF graphs.
[0136] Construct multiple SARIMA models, such as SARIMA(1,1,1)(1,1,1,12), SARIMA(2,1,1)(1,1,1,12), SARIMA(1,1,2)(1,1,1,12), SARIMA(2,1,2)(1,1,1,12).
[0137] Fit each SARIMA model and calculate the AIC and BIC values. Assume that the minimum AIC value of SARIMA(1,1,1)(1,1,1,12) is 100. Select SARIMA(1,1,1)(1,1,1,12) as the best model.
[0138] In this embodiment, the calculation formula of the AIC value is:
[0139]
[0140] Among them, Likelihood(L) is the likelihood of historical toll data; K is the number of parameters of the model; ∫a(x)*f(x)dx is the fitting error of toll data; ∑[γ j *g(t j ) represents the contribution of a preset road section or toll booth; k The weight of the preset road section or toll booth; represents the probability contribution of a preset road section or toll station; k is the traffic volume of the preset road section or toll station; h(z k ) is z k The service level of k ) is z k The probability density function of γ j is the weight of the preset road section or toll station j; g(t j ) is t j The degree of traffic congestion; j is the characteristic variable of road section or toll station j; α(x) is the weight function of toll data x; f(x) is the normal distribution function of toll data x;
[0141] The calculation formula of the BIC value is:
[0142] BIC=-2×log(∫α(x)f(x)dx)+log(h(z k ))
[0143] Where log(∫α(x)f(x)dx) is the fitting error of the charging data.
[0144] According to the scheme provided by the present invention, it includes: step S1, identifying and matching the vehicle travel path according to the license plate information, the highway network topology information and the geographical location information of the toll station to obtain the vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set and a disputed unmatched path set; step S2, according to the actual mileage of the vehicle, the differentiated charging standards of each section and the vehicle path information, the tolls in the undisputed path set and the disputed matched path set are split to obtain the first split fee; wherein, for vehicles traveling on multiple paths in the road network, based on the historical flow distribution and the congestion of adjacent paths, Information, predict the passing probability of the vehicle on the vehicle path information according to the Markov path selection model, adjust the first split fee to obtain the second split fee data according to the passing probability, the road section service level index and the difference in the distance of the alternative path; step S3, target optimization of the second split fee data according to the gradient descent hybrid algorithm to obtain the third split fee data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data; step S4, calculate the error index according to the third split fee data and the actual charging data, and predict the toll revenue and total revenue of each year of the whole life cycle according to the error index and the historical charging data. The present invention uses a big data processing method to accurately split multi-path tolls and predict the toll revenue of the whole life cycle, with low error and high accuracy, exploring new ways for project investment benefit measurement.
[0145] Figure 3 The schematic diagram of the framework of the multi-path toll splitting device based on big data processing according to an embodiment of the present invention is shown. The multi-path toll splitting device based on big data processing includes:
[0146] The path identification module 310 is used to identify and match the vehicle travel path according to the license plate information, the highway network topology information and the toll station geographical location information to obtain the vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set and a disputed unmatched path set;
[0147] The path splitting module 320 is used to split the tolls of the paths in the undisputed path set and the disputed matched path set according to the actual mileage of the vehicle, the differentiated charging standards of each road section and the vehicle path information to obtain a first split fee; wherein, for vehicles passing through multiple paths in the road network, based on the historical traffic distribution and the congestion information of the adjacent paths, the passing probability of the vehicle on the vehicle path information is predicted according to the Markov path selection model, and the passing probability, the road section service level index and the difference in the distance of the alternative paths are adjusted to obtain the second split fee data;
[0148] A target optimization module 330 is used to perform target optimization on the second split cost data according to a gradient descent hybrid algorithm to obtain third split cost data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data;
[0149] The error evaluation module 340 is used to calculate the error index based on the third split cost data and the actual charging data, and predict the charging income and total income of each year of the whole life cycle based on the error index and the historical charging data.
[0150] Figure 4 The schematic diagram of the structure of the computing device embodiment of the present invention is shown. The specific embodiment of the present invention does not limit the specific implementation of the computing device.
[0151] like Figure 4 As shown, the computing device may include: a processor (processor) 402 , a communications interface (Communications Interface) 404 , a memory (memory) 406 , and a communication bus 408 .
[0152] The processor 402, the communication interface 404, and the memory 406 communicate with each other via the communication bus 408. The communication interface 404 is used to communicate with other devices such as a client or other server network elements. The processor 402 is used to execute the program 410, which can specifically execute the relevant steps in the above-mentioned multi-path toll splitting method embodiment based on big data processing.
[0153] Specifically, the program 410 may include program codes, which include computer operation instructions.
[0154] The processor 402 may be a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the computing device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0155] The memory 406 is used to store the program 410. The memory 406 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0156] According to the scheme provided by the present invention, it includes: step S1, identifying and matching the vehicle travel path according to the license plate information, the highway network topology information and the geographical location information of the toll station to obtain the vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set and a disputed unmatched path set; step S2, according to the actual mileage of the vehicle, the differentiated charging standards of each section and the vehicle path information, the tolls in the undisputed path set and the disputed matched path set are split to obtain the first split fee; wherein, for vehicles traveling on multiple paths in the road network, based on the historical flow distribution and the congestion of adjacent paths, Information, predict the passing probability of the vehicle on the vehicle path information according to the Markov path selection model, adjust the first split fee to obtain the second split fee data according to the passing probability, the road section service level index and the difference in the distance of the alternative path; step S3, target optimization of the second split fee data according to the gradient descent hybrid algorithm to obtain the third split fee data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data; step S4, calculate the error index according to the third split fee data and the actual charging data, and predict the toll revenue and total revenue of each year of the whole life cycle according to the error index and the historical charging data. The present invention uses a big data processing method to accurately split multi-path tolls and predict the toll revenue of the whole life cycle, with low error and high accuracy, exploring new ways for project investment benefit measurement.
[0157] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and set in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition they can be divided into multiple submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner can be combined in any combination. Unless otherwise explicitly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose. In addition, those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments means being within the scope of the present invention and forming different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination. The present invention can be implemented by hardware including several different elements and by appropriately programmed computers. In a unit claim that lists several means, several of these means may be embodied by the same hardware item. The steps in the above embodiments should not be understood as limiting the execution order unless otherwise specified.
Claims
1. A multi-path toll splitting method based on big data processing, characterized in that: include: Step S1, identifying and matching vehicle travel paths according to license plate information, highway network topology information, and toll station geographic location information to obtain vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set, and a disputed unmatched path set; Step S2, according to the actual mileage of the vehicle, the differentiated charging standards of each road section and the vehicle path information, the tolls of the paths in the undisputed path set and the disputed matched path set are split to obtain a first split fee; wherein, for vehicles passing through multiple paths in the road network, based on the historical traffic distribution and the congestion information of the adjacent paths, the passing probability of the vehicle on the vehicle path information is predicted according to the Markov path selection model, and according to the passing probability, the road section service level index and the difference in the distance of the alternative paths, the first split fee is adjusted to obtain the second split fee data; Step S3, performing target optimization on the second split cost data according to a gradient descent hybrid algorithm to obtain third split cost data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data; Step S4, calculating an error index based on the third split cost data and the actual charging data, and predicting the charging income and total income for each year of the whole life cycle based on the error index and the historical charging data.
2. The multi-path toll splitting method based on big data processing according to claim 1 is characterized by: The license plate information includes the license plate number, vehicle type, vehicle axle number and vehicle exhaust emission standard; The highway network topology information includes the section number, section length, section connection relationship, section type, section design speed, section slope and section curvature; The toll station geographic location information includes the toll station number, the toll station geographic coordinates, the toll station service range, the number of toll station lanes and the toll station ETC / MTC ratio.
3. The multi-path toll splitting method based on big data processing according to claim 1 is characterized in that: The adjusting the first split cost to obtain the second split cost data further comprises: Step S21, training is performed based on the historical traffic data of the vehicle, and the state transfer matrix and emission matrix of the HMM are constructed by using the maximum likelihood estimation method; wherein the state space is a set of road sections; the state transfer probability is determined based on the historical traffic data and the connectivity of the road sections, indicating the probability of a vehicle transferring from one road section to another; the observation space is the multimodal perception data of the vehicle, which is used to infer the current state of the vehicle; the emission probability is calculated based on the matching degree between the vehicle position and the road section; Step S22, dynamically adjusting the state transition probability of the HMM according to the congestion level of the adjacent paths; Step S23, decoding the HMM according to the Viterbi algorithm to predict the passing probability of the vehicle on each possible path; Step S24, adjusting the first split cost by a correction factor according to the road section service level index and the difference in the alternative path distance to obtain second split cost data.
4. The multi-path toll splitting method based on big data processing according to claim 1 is characterized in that: The performing target optimization on the second split cost data according to the gradient descent hybrid algorithm to obtain the third split cost data further comprises: Step S31, defining an objective function for measuring the error between the split fee and the actual charging data; Step S32, calculating the gradient of the objective function with respect to the splitting cost, wherein the gradient represents the change direction and magnitude of the objective function under the current splitting cost; Step S33, updating the first-order moment estimate and the second-order moment estimate according to the Adam algorithm; wherein the first-order moment estimate is the average exponential shift of the gradient, and the second-order moment estimate is the average exponential shift of the gradient square; Step S34, updating the splitting cost data according to the update rule of the Adam algorithm until the objective function converges or reaches a preset number of iterations, and outputting the optimized third splitting cost data.
5. The multi-path toll splitting method based on big data processing according to claim 1 is characterized in that: The prediction of charging income and total income for each year of the life cycle based on the error index and historical charging data further includes: Step S51, model the historical toll data according to the seasonal SARIMA model, and solve the model order (p, d, q) or (p, d, q)(P, D, Q) s , so that the AIC or BIC value of the seasonal SARIMA model is minimized; Step S52: predicting the total income of the entire life cycle according to the seasonal SARIMA model.
6. The multi-path toll splitting method based on big data processing according to claim 1 is characterized in that: The calculation formula of the alternative path distance difference is: Among them, d 1i and d 2i are the lengths of the i-th type of road on the current path and the alternative path respectively; w i is the weight of the i-th type of road.
7. The multi-path toll splitting method based on big data processing according to claim 5 is characterized in that: The order of the solution model is (p, d, q) or (p, d, q)(P, D, Q) s , minimizing the AIC or BIC value of the seasonal SARIMA model further includes: Repeat the first-order difference of the historical toll data, record the order d of the difference, and repeat the seasonal difference of the historical toll data, record the order D of the seasonal difference; According to the autocorrelation function and the partial autocorrelation function, the AR and MA orders (p, q, P, Q) are preliminarily estimated, and multiple SARIMA models are constructed according to the orders (p, q, P, Q); Fit various SARIMA models according to the historical charging data, calculate the AIC value and BIC value of each SARIMA model, and select the SARIMA model with the smallest AIC or BIC value as the best model.
8. The multi-path toll splitting method based on big data processing according to claim 7 is characterized in that: The calculation formula of the AIC value is: Among them, Likelihood(L) is the likelihood of historical toll data; K is the number of parameters of the model; ∫α(x)*f(x)dx is the fitting error of toll data; ∑[γ j *g(t j ) represents the contribution of a preset road section or toll booth; k The weight of the preset road section or toll booth; represents the probability contribution of a preset road section or toll station; k is the traffic volume of the preset road section or toll station; h(z k ) is z k The service level of k ) is z k The probability density function of γ j is the weight of the preset road section or toll station j; g(t j ) is t j The degree of traffic congestion; j is the characteristic variable of road section or toll station j; α(x) is the weight function of toll data x; f(x) is the normal distribution function of toll data x; The calculation formula of the BIC value is: BIC=-2×log(∫α(x)f(x)dx)+log(h(z k )) Where log(∫α(x)f(x)dx) is the fitting error of the charging data.
9. A multi-path toll splitting device based on big data processing, characterized in that: The method for splitting multi-path tolls according to any one of claims 1 to 8 comprises: A path identification module is used to identify and match the vehicle travel path according to the license plate information, the highway network topology information and the toll station geographical location information to obtain the vehicle path information, wherein the vehicle path information includes an undisputed path set, a disputed matched path set and a disputed unmatched path set; A path splitting module is used to split the tolls of the paths in the undisputed path set and the disputed matched path set according to the actual mileage of the vehicle, the differentiated charging standards of each road section and the vehicle path information to obtain a first split fee; wherein, for vehicles passing through multiple paths in the road network, based on the historical traffic distribution and the congestion information of the adjacent paths, the passing probability of the vehicle on the vehicle path information is predicted according to the Markov path selection model, and the first split fee is adjusted according to the passing probability, the road section service level index and the difference in the distance of the alternative paths to obtain the second split fee data; A target optimization module, used for performing target optimization on the second split cost data according to a gradient descent hybrid algorithm to obtain third split cost data, wherein the objective function of the target optimization is to minimize the error between the split result and the actual charging data; The error evaluation module is used to calculate the error index based on the third split cost data and the actual charging data, and predict the charging income and total income of each year of the whole life cycle based on the error index and historical charging data.
10. A computing device comprising: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the multi-path toll splitting method based on big data processing as described in any one of claims 1-8.