A highway section mileage measurement method based on ETC big data

By constructing a highway section mileage generation model using ETC big data, the problems of resource consumption and safety hazards of traditional measurement methods have been solved, achieving efficient and safe mileage measurement and enhancing the management and application value of the ETC system.

CN115938105BActive Publication Date: 2025-11-11FUJIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210729479.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-11-11
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

Traditional methods of measuring highway mileage consume a lot of manpower, material resources, and financial resources, and pose safety risks, making it difficult to achieve efficient and secure ETC network toll collection.

Method used

A highway section mileage measurement method based on ETC big data is adopted. By acquiring geographical coordinates, travel time and dwell time, a section mileage generation model is constructed using Gaussian random vectors and normal distribution models. Combined with box plots to clean up noisy data, accurate mileage measurement is achieved.

Benefits of technology

It greatly reduces the cost of manual surveying, improves the safety and accuracy of measurements, enhances the basic information of the ETC system, and contributes to the refined management of highways and the application value of ETC big data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115938105B_ABST
    Figure CN115938105B_ABST
Patent Text Reader

Abstract

The application discloses a highway section mileage measuring method based on ETC big data, comprising the following steps: acquiring geographical position coordinates of a starting point and an ending point of a highway section, and acquiring total mileage of the section by using a map API; acquiring driving time of a vehicle in the whole section, and calculating average driving speed of the vehicle in the whole section in combination with the mileage of the whole section; dividing the whole section into several sections, acquiring residence time of the vehicle in each section, and calculating section driving mileage by taking the average speed of the whole section as driving speed of each section; and constructing a section mileage generation model according to different section driving mileages of the vehicle. The model constructed by the application realizes measurement of highway section mileage only by relying on ETC big data, has small measurement error and stable performance, and is helpful to fine management of highways and improvement of application value of ETC big data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of highway management technology, and in particular to a method for measuring highway section mileage based on ETC big data. Background Technology

[0002] By the end of 2020, my country's total expressway mileage reached 161,000 kilometers, ranking first in the world. To further improve the operational efficiency of my country's expressways, by the end of 2019, my country's Electronic Toll Collection (ETC) system had been connected in 29 provinces across the country, with 24,588 ETC gantry systems built and 48,211 ETC lanes upgraded. The cumulative number of ETC users nationwide exceeded 200 million.

[0003] To ensure accurate and reasonable ETC (Electronic Toll Collection) network tolling, according to my country's regulations for newly built expressways, the mileage of a new expressway must be measured on-site before ETC network tolling is implemented to determine the actual toll mileage and calculate the toll amount accordingly. However, given my country's vast and complex expressway network, traditional measurement methods not only consume significant manpower, material resources, and financial resources, but also pose safety hazards during on-site measurements. Summary of the Invention

[0004] The purpose of this invention is to provide a method for measuring highway section mileage based on ETC big data, which greatly reduces the cost of manual surveying and improves the safety factor.

[0005] The technical solution adopted in this invention is:

[0006] A method for measuring highway section mileage based on ETC big data includes the following steps:

[0007] Step 1: Obtain the geographical coordinates of the start and end points of the highway segment, and use the map API to obtain the total mileage of the segment;

[0008] Step 2: Obtain the vehicle's travel time over the entire road segment and calculate the vehicle's average speed over the entire road segment.

[0009] Step 3: Obtain the vehicle's dwell time in each section, and calculate the mileage of each section using the average speed of the road segment as the driving speed of each section.

[0010] Specifically, first obtain the mileage of the entire road segment according to step 1, then obtain the vehicle's travel time on the entire road segment according to vehicle data, and calculate the travel speed.

[0011] Then, the average speed of the entire segment is calculated, the dwell time of the vehicle in each segment is obtained, the average speed is used as the segment speed, and then the segment mileage of each segment is obtained.

[0012] Step 4: Based on the vehicle's segmented mileage, construct the segmented mileage generation model as follows:

[0013] ΔD~N(μ,Γ Δd );

[0014] Where ΔD is a Gaussian random vector with mean vector μ = [μ1, μ2, ..., μ... n-1 ], μ n The mean distance of the nth segment is calculated based on the mileage traveled by the vehicles passing through it; the covariance matrix is... This represents the variance of each segment, where m represents the number of vehicles passing through the segment, and N represents a normal distribution. In other words, it calculates the mileage of each vehicle in different segments based on the vehicles passing through, inputs this data into the model to obtain the total mileage of the entire road segment and the mileage of each segment.

[0015] Furthermore, in step 1, the map API is the Gaode Map API.

[0016] Furthermore, in step 2, the average speed of the vehicle is calculated using the average speed formula:

[0017]

[0018] Where d is the total mileage of the road segment. It indicates the total travel time for the entire route.

[0019] Furthermore, in step 3, a noise data cleaning method based on box plots is used to obtain the vehicle's dwell time and mileage in each segment.

[0020] Furthermore, the specific steps of step 3 are as follows:

[0021] Step 301: Select road segment LD where the traffic conditions are free-flowing, and the number of lanes in each section of the road segment is the same; if the traffic conditions of the road segment do not meet the requirements, select a time period such as late night when the traffic conditions meet the requirements; if the number of lanes in each section is different, divide the road segment into multiple sections and process them independently.

[0022] Step 302: Due to variations in vehicle speed and driving habits, the proportion of a vehicle's dwell time in a certain section relative to the total travel time of the entire section is called the section time ratio r.

[0023]

[0024] Where Δt is the dwell time of the segment, The total travel time for the entire route.

[0025] Step 303: Normalize the dwell time of the section to obtain the proportion R of the time taken by m vehicles in the entire road segment;

[0026]

[0027] Where m refers to m vehicles; n refers to the nth segment;

[0028] Step 304: Clean the noise data generated by vehicles passing through the service area using a box plot to obtain the effective vehicle subset M” of the road segment:

[0029]

[0030] The effective subset of vehicles is the set of vehicles on the road segment used.

[0031] Step 305, obtain the segment mileage ΔD:

[0032]

[0033] Where, V = diag(v 1 ,v 2 ,…,v n-1 ) represents the speed of each vehicle in the subset; This represents the mileage traveled by the m-th vehicle in the (n-1)th segment; m represents the number of vehicles in the effective vehicle subset M”.

[0034] Step 3041: When a vehicle passes through a service area, the time spent is divided into those who stop at the service area and those who do not. Different time percentages for each segment can be obtained. The time percentage for vehicles that do not stop at service areas is:

[0035]

[0036] Where Δt1 is the dwell time in this section when not stopping, and Δt is the dwell time of the entire road segment including this section;

[0037] The percentage of time spent on each section by vehicles stopping at service areas is as follows:

[0038]

[0039] Where Δt1 is the dwell time of the vehicle in the section, which is not included in the service area time; Δt s This refers to the time spent at the service area.

[0040] Step 3042: Since vehicle stops at service areas cause the segment travel time ratio to deviate from the overall true distribution, box plots are needed to clean the data. The required data for the box plot include: Q1 (first quartile); Q2 (second quartile, also known as the median); Q3 (third quartile); IQR = Q3 - Q1 (interquartile range); Q1 - 1.5 × IQR (lower limit) and Q3 + 1.5 × IQR (upper limit); Outliers are noise points, whose values ​​are greater than Q3 + 1.5 × IQR or less than Q1 - 1.5 × IQR, also known as outliers or anomalies.

[0041] Step 3043: Construct the vehicle subset I of the stopover service area based on the box plot results. j :

[0042]

[0043] Where J = {j} (j∈[1,n-1]) is the set of segments where the service area is located; Let i be the proportion of time taken by vehicle i in segment j; The third quartile of the time taken for segment j in the box plot; The interquartile range of the time taken for segment j in the box plot; The first quartile of the time taken for segment j in the box plot;

[0044] Step 3044, construct the vehicle subset M' of the road segment:

[0045]

[0046] Where M is the original vehicle set, and M' is the vehicle set after removing vehicles from the stop service area;

[0047] Step 3045: Due to the complex and ever-changing traffic conditions, including traffic flow, road maintenance, and unforeseen events, M' needs further cleaning, and then a subset I' of abnormal vehicles in the section needs to be constructed. j :

[0048]

[0049] in, The percentage of time taken for abnormal vehicles in a given section; This is the third quartile in the box plot; This represents the interquartile range in the box plot; This is the first quartile in the box plot;

[0050] Step 3046, thus obtaining the effective vehicle subset M of the road segment:

[0051]

[0052] The effective subset of vehicles is the set of vehicles on the road segment in which they are used.

[0053] Furthermore, the steps in step 4 for constructing the gantry section mileage generation model using the law of large numbers are as follows:

[0054] Step 401: Based on the segment mileage obtained in Step 3, it can be seen that the mileage of each vehicle in different segments is independent of each other and follows the same distribution with mathematical expectation.

[0055]

[0056] in, μ represents the mileage traveled by the i-th vehicle in segment j. j Let be the mathematical expectation of the mileage of road segment j.

[0057] Step 402, according to the law of large numbers, the sequence is... Converges in probability to μ j ,Right now set up With variance According to the central limit theorem, Δd j sum Standardized variable Y m for:

[0058]

[0059] Among them, Y m The distribution function F m (x) satisfies for any x:

[0060]

[0061] Among them, F m Φ(x) is the distribution function, and Φ(x) is the standard normal distribution function.

[0062] Step 403, when m is large enough (m≥30) Δd j After appropriate standardization, the mean of the distribution converges to a normal distribution. Therefore, any... mean Approximately follows the mean μ j The variance is The mileage follows a normal distribution. Therefore, a gantry section mileage generation model (MGM) based on ETC big data can be constructed:

[0063] ΔD~N(μ,Γ Δd )

[0064] Where ΔD is the generated segment mileage, and its mean vector μ = [μ1, μ2, ..., μ n-1 ], μ n The mean distance of the nth segment is calculated based on the mileage traveled by the vehicles passing through it; the covariance matrix is... Let represent the variance of each segment, and m represent the m vehicles that passed through the segment.

[0065] This invention adopts the above technical solution and relies on ETC big data to realize the measurement of section mileage. The measurement accuracy is high and the performance is stable. It changes the traditional highway mileage measurement operation mode, which not only greatly reduces the cost of manual surveying and mapping, but also improves the safety factor. At the same time, it improves the basic information of the ETC system, which helps to refine the management of highways and enhance the application value of ETC big data. Attached Figure Description

[0066] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments;

[0067] Figure 1 This is a schematic diagram of the structure of a highway section mileage measurement method based on ETC big data according to the present invention;

[0068] Figure 2 This is a schematic diagram of a noise data cleaning model for box plots.

[0069] Figure 3 Schematic diagram of LD1;

[0070] Figure 4 Schematic diagram of LD2;

[0071] Figure 5 A diagram showing the distribution of mileage data before and after LD1 data cleaning;

[0072] Figure 6 A diagram showing the relative error of segment mileage before and after LD1 data cleaning;

[0073] Figure 7 A diagram showing the distribution of mileage data before and after LD2 data cleaning;

[0074] Figure 8 A diagram showing the relative error of segment mileage before and after LD2 data cleaning;

[0075] Figure 9 This is a schematic diagram of the overall mileage error of LD1;

[0076] Figure 10 This is a schematic diagram of the overall mileage error of LD2;

[0077] Figure 11This is a schematic diagram illustrating the cumulative error probability of a single segment in LD1.

[0078] Figure 12 This is a schematic diagram illustrating the cumulative error probability of a single segment in LD2.

[0079] Figure 13 This is a schematic diagram illustrating the impact of different parameters on model error in LD1;

[0080] Figure 14 This is a schematic diagram illustrating the impact of different parameters on model error in LD2. Detailed Implementation

[0081] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0082] like Figures 1 to 14 As shown in the figure, this invention discloses a method for measuring highway section mileage based on ETC big data, which includes the following steps:

[0083] Step 1: Obtain the geographical coordinates of the start and end points of the highway segment, and use the map API to obtain the total mileage of the segment;

[0084] Step 2: Obtain the vehicle's travel time over the entire road segment and calculate the vehicle's average speed over the entire road segment.

[0085] Step 3: Obtain the vehicle's dwell time in each section, and calculate the mileage of each section using the average speed of the road section as the driving speed of each section.

[0086] Step 4: Based on the vehicle's segmented mileage, construct the segmented mileage generation model as follows:

[0087] ΔD~N(μ,Γ Δd )

[0088] Where ΔD is a Gaussian random vector with mean vector μ = [μ1, μ2, ..., μ... n-1 ], μ n The mean distance of the nth segment is calculated based on the mileage traveled by the vehicles passing through it; the covariance matrix is... Let represent the variance of each segment, and m represent the m vehicles that passed through the segment.

[0089] Furthermore, in step 1, the map API is the Gaode Map API.

[0090] Furthermore, in step 2, the average speed of the vehicle is calculated using the average speed formula:

[0091]

[0092] Where d is the total mileage of the road segment. This indicates the travel time for the entire route.

[0093] Furthermore, in step 3, a noise data cleaning method based on box plots is used to obtain the vehicle's dwell time and mileage in each segment.

[0094] Furthermore, the specific steps of step 3 are as follows:

[0095] Step 301: Select road segment LD where the traffic conditions are free-flowing, and the number of lanes in each section of the road segment is the same; if the traffic conditions of the road segment do not meet the requirements, select a time period such as late night when the traffic conditions meet the requirements; if the number of lanes in each section is different, divide the road segment into multiple sections and process them independently.

[0096] Step 302: Due to variations in vehicle speed and driving habits, the proportion of a vehicle's dwell time in a certain section relative to the total travel time of the entire section is called the section time ratio r.

[0097]

[0098] Where Δt is the dwell time of the segment, This refers to the travel time for the entire route.

[0099] Step 303: Normalize the dwell time of the section to obtain the proportion R of the time taken by m vehicles in the entire road segment;

[0100]

[0101] Where m refers to m vehicles; n refers to the nth segment;

[0102] Step 304: Clean the noise data generated by vehicles passing through the service area using a box plot to obtain the effective vehicle subset M” of the road segment:

[0103]

[0104] The effective subset of vehicles is the set of vehicles on the road segment used.

[0105] Step 305, obtain the segment mileage ΔD:

[0106]

[0107] Where, V = diag(v 1 ,v 2 ,…,v n-1 ) represents the speed of each vehicle in the subset; This represents the mileage traveled by the m-th vehicle in the (n-1)th segment; m represents the number of vehicles in the effective vehicle subset M”.

[0108] Furthermore, the specific steps of step 304 are as follows:

[0109] Step 3041: When a vehicle passes through a service area, the time spent is divided into those who stop at the service area and those who do not. Different time percentages for each segment can be obtained. The time percentage for vehicles that do not stop at service areas is:

[0110]

[0111] Where Δt1 is the dwell time in this section when not stopping, and Δt is the dwell time of the entire road segment including this section;

[0112] The percentage of time spent on each section by vehicles stopping at service areas is as follows:

[0113]

[0114] Where Δt1 is the dwell time of the vehicle in the section, which is not included in the service area time; Δt s This refers to the time spent at the service area.

[0115] Step 3042: Since vehicle stops at service areas cause the segment travel time ratio to deviate from the overall true distribution, box plots are needed to clean the data. The required data for the box plot include: Q1 (first quartile); Q2 (second quartile, also known as the median); Q3 (third quartile); IQR = Q3 - Q1 (interquartile range); Q1 - 1.5 × IQR (lower limit) and Q3 + 1.5 × IQR (upper limit); Outliers are noise points, whose values ​​are greater than Q3 + 1.5 × IQR or less than Q1 - 1.5 × IQR, also known as outliers or anomalies.

[0116] Step 3043: Construct the vehicle subset I of the stopover service area based on the box plot results. j :

[0117]

[0118] Where J = {j} (j∈[1,n-1]) is the set of segments where the service area is located; Let i be the proportion of time taken by vehicle i in segment j; The third quartile of the time taken for segment j in the box plot; The interquartile range of the time taken for segment j in the box plot; The first quartile of the time taken for segment j in the box plot;

[0119] Step 3044, construct the vehicle subset M' of the road segment:

[0120]

[0121] Where M is the original vehicle set, and M' is the vehicle set after removing vehicles from the stop service area;

[0122] Step 3045: Due to the complex and ever-changing traffic conditions, including traffic flow, road maintenance, and unforeseen events, M' needs further cleaning, and then a subset I' of abnormal vehicles in the section needs to be constructed. j :

[0123]

[0124] in, The percentage of time taken for abnormal vehicles in a given section; This is the third quartile in the box plot; The interquartile interval in the box plot; This is the first quartile in the box plot;

[0125] Step 3046, thus obtaining the effective vehicle subset M of the road segment:

[0126]

[0127] The effective subset of vehicles is the set of vehicles on the road segment used.

[0128] Furthermore, the steps in step 4 for constructing the gantry section mileage generation model using the law of large numbers are as follows:

[0129] Step 401: Based on the segment mileage obtained in Step 3, it can be seen that the mileage of each vehicle in different segments is independent of each other and follows the same distribution with mathematical expectation.

[0130]

[0131] in, μ represents the mileage traveled by the i-th vehicle in segment j. j Let be the mathematical expectation of the mileage of road segment j.

[0132] Step 402, according to the law of large numbers, the sequence is... Converges in probability to μ j ,Right now set up With variance According to the central limit theorem, Δd j sum Standardized variable Y m for:

[0133]

[0134] Among them, Y m The distribution function F m (x) satisfies for any x:

[0135]

[0136] Among them, F m Φ(x) is the distribution function, and Φ(x) is the standard normal distribution function.

[0137] Step 403, when m is large enough (m≥30) Δd j After proper standardization, the mean of the distribution converges to a normal distribution. Therefore, any... mean Approximately follows the mean μ j The variance is The mileage follows a normal distribution. Therefore, a gantry section mileage generation model (MGM) based on ETC big data can be constructed:

[0138] ΔD~N(μ,Γ Δd )

[0139] Where ΔD is the generated segment mileage, and its mean vector μ = [μ1, μ2, ..., μ n-1 ], μ n The mean distance of the nth segment is calculated based on the mileage traveled by the vehicles passing through it; the covariance matrix is... Let represent the variance of each segment, and m represent the m vehicles that passed through the segment.

[0140] The specific principles of this invention will be explained in detail below:

[0141] Assume that the geographical coordinates (latitude and longitude) of the starting and ending nodes of road segment LD are known. If a road segment LD does not meet assumption 1, then nodes with clearly defined geographical coordinates, such as tollbooth entrances and exits, can be selected as the starting and ending nodes of the road segment. Obviously, this assumption is reasonable.

[0142] Therefore, based on the geographical coordinates of the start and end points of road segment LD, the total mileage of road segment LD can be obtained as d through the Gaode Map driving route planning API. Then, the average driving speed v of the vehicle throughout the entire road segment is:

[0143]

[0144] However, factors such as the number of lanes on a road segment, unexpected events, and driving behavior inevitably affect vehicle speed. To achieve a more ideal result and meet the requirements of the MGM model, the following constraints are imposed on the study road segment:

[0145] L1 = L2 = ... = Ln-1, where Lj is the number of lanes in segment j;

[0146] The selected road segment LD is in free-flow traffic condition.

[0147] If a road segment does not meet constraint 1, it can be divided into multiple segments and processed independently; if the road conditions of a road segment do not meet constraint 2, a time period that meets the free flow of traffic, such as late at night, can be selected. Therefore, this constraint is reasonable.

[0148] Assuming all vehicles operate under ideal traffic conditions, with no interference or influence between them, and vehicles having free-flow speed.

[0149] Based on the word hypothesis, it can be clearly assumed that They are independent, follow the same distribution, and have mathematical expectation. According to the law of large numbers, the sequence... Converges in probability to μ j ,Right now Let's assume With variance According to the central limit theorem, Δd j sum Standardized variable Y m for:

[0150]

[0151] Among them, Y m The distribution function F m (x) satisfies for any x:

[0152]

[0153] From the above formula, it can be seen that when m is large enough (usually requiring a large sample size m≥30), Δd j After proper standardization, the mean of the distribution converges to a normal distribution. Therefore, any... mean Approximately follows the mean μ j The variance is The mileage follows a normal distribution. Therefore, a gantry section mileage generation model (MGM) based on ETC big data can be constructed:

[0154] ΔD~N(μ,Γ Δd )

[0155] ΔD is a Gaussian random vector with mean vector μ = [μ1, μ2, ..., μ]. n-1 and covariance matrix

[0156] Table 1: Description of Some Fields in ETC Transaction Data

[0157]

[0158]

[0159] Table 1 shows the descriptions of some fields in the ETC transaction data. After the data cleaning and generation model were effectively validated, the distribution characteristics of the ETC transaction data of LD1 and LD2 were analyzed. The data distribution before and after cleaning of the segment mileage is shown in Table 1. Figure 5 and Figure 6 Before data cleaning, although the mileage data distribution generally exhibited a Gaussian distribution, a large amount of noise remained, causing a skewed distribution. Specifically, QD1, QD2, and QD3 were all service area segments, with a large number of mileage data points scattered to the left of the actual mileage, resulting in a long tail on the left side of the fitted curve (top), exhibiting a negatively skewed distribution; while QD4 was a service area segment, with a large number of mileage data points scattered to the right of the actual mileage, resulting in a long tail on the right side of the fitted curve (top), exhibiting a positively skewed distribution. Further cleaning of the outlier mileage data was performed using the ODC algorithm. Figure 5 and Figure 6 As can be seen, compared with the horizontal axis coordinates of each segment before cleaning, all vertical axis coordinates have shrunk to a certain range, indicating that all mileage data after cleaning are close to the true mileage; at the same time, from the fitting curves of each segment after cleaning (right side), it can be seen that the mileage data distribution of all segments shows good Gaussian distribution characteristics, indicating that most of the abnormal mileage data has been cleaned well.

[0160] Further mileage data before and after cleaning were used to estimate the mileage of each segment, and MRE was used as the evaluation index. Specifically, 100 groups were randomly sampled from the LD1 and LD2 datasets, with sample sizes ranging from 1 to 200 per group, to study the fluctuation evolution of MRE. The results are as follows: Figure 7 and Figure 8As shown, all segments exhibit the characteristic that the smaller the sample size, the more drastic the MRE fluctuations. Specifically, before data cleaning, the MRE values ​​of all segments fluctuated widely; segments without service areas converged to negative values, while segments with service areas converged to positive values, showing consistent distribution characteristics with those before cleaning, further highlighting the skewness of the data distribution. After cleaning, the MRE values ​​of all segments fluctuated only within a small range, exhibiting slight oscillations followed by rapid convergence to 0. Furthermore, the degree of mileage deviation after final convergence revealed that, regardless of whether it was a service area segment or a segment without service areas, the longer the segment mileage, the greater the deviation. After cleaning, the total number of complete LD1 and LD2 trajectories were obtained, amounting to 1033 and 1302 respectively. A comparative analysis of the data distribution characteristics and MRE values ​​of each segment before and after data cleaning was conducted.

[0161] While ensuring data quality, sample size directly affects the magnitude of the generated model error. To fully investigate the impact of sample size on error, MAE (Model Error Evolution) is used as the evaluation metric. Specifically, random samples are taken from the LD1 and LD2 datasets, with sample sizes ranging from 1 to 1000, and the study is repeated 100 times to investigate the fluctuation and evolution of MAE. Figure 9 and Figure 11 As shown, the MAE decreases rapidly with increasing sample size, and its error fluctuation range gradually decreases, eventually stabilizing. According to the sample size requirements in Section 2, m = 30 is a watershed for sample size. When m < 30, the MAE error fluctuates drastically and over a wide range, with some errors exceeding 200m, indicating poor performance of the generative model under small sample sizes. When m > 30, the MAE fluctuation range is smaller, and its mean fluctuation is gentle and below 50m, indicating a significant improvement in generative model performance under large sample sizes. This further verifies the requirement for large sample sizes in the generative model in Section 2.

[0162] Further experiments were conducted on each section of the road segment individually, and the overall probability distribution of MAE was statistically analyzed. Let δ = 100m and α = 0.02. Since the standard deviation σ is unknown, 10,000 groups were randomly selected from the dataset, with 30 samples per group. The standard deviation of each segment was calculated, and the maximum value was taken: σ1 ≈ [379, 86, 279, 348] (unit: m, the same below) for each segment of LD1, and σ2 ≈ [115, 315, 152, 370] for each segment of LD2. Therefore, let σ1 and σ2 be the standard deviations of LD1 and LD2. max The values ​​are 379m and 370m, respectively. According to Corollary 1, the required sample sizes for LD1 and LD2 are at least 78 and 75, respectively. Therefore, we only need to set m1 = 78 and m2 = 75. To approximate the true error probability distribution for each segment, 10,000 sets of experiments (the same below) are conducted for each segment, and the cumulative error probability distribution (CDF) curve is plotted. For example... Figure 10 and Figure 12As shown, all segments can control the error within 100m with a 100% confidence probability, and the probability value is consistently higher than the preset value of 98%. Meanwhile, at a 98% confidence probability, each segment in LD1 can control the error within 73m, while each segment in LD2 can control the error within 66m, indicating significant performance of the generative model. In particular, with the same sample size, QD2 of LD1 and QD1 of LD2 perform better, controlling the error within 30m and 32m respectively with a 100% confidence probability. Further research revealed that the mileage of both segments is relatively short, indicating better performance of their generative models.

[0163] To further investigate the impact of sample size on the error of the generative model under different parameters, the following experiment was conducted using the controlled variable method:

[0164] To study the effect of parameter σ variation, we set α and δ to fixed values. Let α = 0.02, δ = 100m, and the standard deviations of the two road segments be σ1 and σ2, respectively. We can then obtain the required sample sizes for LD1 and LD2 as [78, 30, 42, 66] and [30, 53, 30, 75], respectively. Their corresponding CDF curves are shown in […]. Figure 13 (a) and (c). Under different σ values, the confidence probabilities of LD1 and LD2 within an error of 100m are [100%, 100%, 99.6%, 99.9%] and [100%, 99.8%, 100%, 100%], respectively, both consistently higher than the preset value of 98%. Meanwhile, at a confidence probability of 98%, LD1 and LD2 can control the error within [64.0, 28.8, 82.9, 77.6] and [31.9, 75.6, 54.3, 62.9], respectively, both significantly lower than the preset value of 100m.

[0165] To study the effect of changes in parameter δ, we will set α and σ to fixed values. Let α = 0.02 and σ = σ max Given δ = [50, 100, 150, 200], the required sample sizes for LD1 and LD2 are [315, 78, 35, 30] and [297, 75, 33, 30], respectively. This indicates that the larger δ is, the smaller the required sample size. However, once δ > 150, the required sample size remains essentially unchanged due to the limitation of a large sample size. Based on the experimental parameter settings, the CDF curves corresponding to the two road segments are shown below. Figure 13 (b) and Figure 14(e) As the sample size increases, the error decreases at the same confidence probability. At a confidence probability of 98%, the errors corresponding to LD1 and LD2 are [25.7, 41.3, 58.0, 62.0] and [22.1, 36.9, 54.3, 56.1], respectively, both significantly smaller than their corresponding preset values. Simultaneously, the confidence probability at the preset value δ reaches 100%, demonstrating a significant effect. Furthermore, as δ decreases, the required sample size increases quadratically, and the steepness of the probability distribution curve increases sequentially, indicating that changes in parameter δ have a significant impact on the error of the generative model.

[0166] To study the effect of changes in parameter α, let δ and σ be fixed values. Without loss of generality, let δ = 100m and σ = σ max Given α = [0.02, 0.04, 0.06, 0.08, 0.10], the required sample sizes for LD1 and LD2 are [78, 61, 51, 44, 39] and [75, 58, 48, 42, 38], respectively. Their CDF curves are shown in […]. Figure 13 (c) and Figure 14 (f). When the confidence probability (1-α) is set, the errors corresponding to LD1 and LD2 are basically distributed between 33-44m and 28-38m, respectively, which are significantly lower than the set threshold of 100m, and the effect is significant. Further research found that the sample size distribution range under different α is small, which makes their probability distribution curves basically consistent, indicating that the change of parameter α has little impact on the error of the generative model.

[0167] This invention employs the above technical solutions to construct a highway segment mileage generation model based on ETC transaction data, and derives the functional relationship between data sample size and model error. Field experimental results show that the model generates segment mileage with an average error of 10m, demonstrating strong robustness and applicability. Finally, ETC transaction data from September 3rd to 5th, 2020, was used for field verification on two sections of the Shenhai Expressway (Caopuyuan Interchange to Neikeng Interchange, and Ganghou Interchange to Xipu Interchange). Theoretical analysis and experimental results show that the model generates segment mileage with an average error of 10m, demonstrating strong robustness and applicability. The model constructed in this invention achieves highway segment mileage measurement solely based on ETC big data, exhibiting small measurement errors and stable performance, which contributes to refined highway management and enhances the application value of ETC big data.

[0168] Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. Without conflict, the embodiments and features in the embodiments of this application can be combined with each other. The components of the embodiments of this application described and illustrated herein can generally be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of this application is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

Claims

1. A method for measuring highway section mileage based on ETC big data, characterized in that: It includes the following steps: Step 1: Obtain the geographical coordinates of the start and end points of the highway segment, and use the map API to obtain the total mileage of the segment; Step 2: Obtain the vehicle's travel time over the entire road segment, and calculate the vehicle's average speed over the entire road segment by combining the total mileage of the entire road segment. The average speed of a vehicle is calculated using the average speed formula: Where d is the total mileage of the road segment. Indicates the total travel time for the entire route; Step 3: Divide the entire road segment into several sections, obtain the dwell time of the vehicle in each section, and calculate the mileage of each section using the average speed of the entire road segment as the driving speed of each section. The specific steps of obtaining the dwell time and mileage of the vehicle in each section using the box plot-based noise data cleaning method in Step 3 are as follows: Step 301: Select road segment LD where the traffic conditions are free-flowing, and the number of lanes in each section of the road segment is the same; if the traffic conditions of the road segment do not meet the requirements, select a time period such as late night when the traffic conditions meet the requirements; if the number of lanes in each section is different, divide the road segment into multiple sections and process them independently. Step 302: Due to variations in vehicle speed and driving habits, the proportion of a vehicle's dwell time in a certain section relative to the total travel time of the entire section is called the section time ratio r. Where Δt is the dwell time of the segment, The total travel time for the entire route. Step 303: Normalize the dwell time of the section to obtain the proportion R of the time taken by m vehicles in the entire road segment; Where m refers to m vehicles; n refers to the nth segment; Step 304: Clean the noise data generated by vehicles passing through the service area using a box plot to obtain the effective vehicle subset M" of the road segment: Among them, M′ is the vehicle set after removing vehicles from stopover service areas, I j ′ represents the subset of abnormal vehicles in the section; j∈[1,n-1] is the set of sections where the service area is located; The effective subset of vehicles is the set of vehicles on the road segment used. Step 305, obtain the segment mileage ΔD: Where, V = diag(v 1 ,v 2 ,…,v n-1 ) represents the speed of each vehicle in the subset; This represents the mileage traveled by the m-th vehicle in the (n-1)th segment; m represents the number of vehicles in the effective vehicle subset M". Step 4: Based on the vehicle's mileage in different segments, construct the segment mileage generation model as follows: ΔD~N(μ,Γ Δd ) Where ΔD is a Gaussian random vector with mean vector μ = [μ1, μ2, ..., μ... n-1 ], μ n The mean distance of the nth segment is calculated based on the mileage traveled by the vehicles passing through it; the covariance matrix is... Let m represent the variance of each segment, m represent the number of vehicles passing through the segment, and N represent a normal distribution.

2. The method for measuring highway section mileage based on ETC big data according to claim 1, characterized in that: In step 1, the map API is the Gaode Map API.

3. The method for measuring highway section mileage based on ETC big data according to claim 1, characterized in that: The specific steps of step 304 are as follows: Step 3041: When a vehicle passes through a service area, the time spent is divided into two categories: when the vehicle stops at the service area and when it does not. Different time percentages are obtained for each segment. The time percentage for vehicles that do not stop at service areas is as follows: Where Δt1 is the dwell time in this section when not stopping, and Δt is the dwell time of the entire road segment including this section; The percentage of time spent on each section by vehicles stopping at service areas is as follows: Where Δt1 is the dwell time of the vehicle in the section, which is not included in the service area time; Δt s This refers to the time spent at the service area. Step 3042, The data needed to obtain the box plot include: Q1 as the first quartile; Q2 as the second quartile, also known as the median; Q3 as the third quartile; IQR = Q3 - Q1 as the interquartile range; Q1 - 1.5 × IQR as the lower limit and Q3 + 1.5 × IQR as the upper limit; and outliers as noise points, whose values ​​are greater than Q3 + 1.5 × IQR or less than Q1 - 1.5 × IQR, also known as outliers or anomalies. Step 3043: Construct the vehicle subset I of the stopover service area based on the box plot results. j : Where I = {j} (j∈[1,n-1]) is the set of segments where the service area is located; Let i be the proportion of time taken by vehicle i in segment j; The third quartile of the time taken for segment j in the box plot; The interquartile range of the time taken for segment j in the box plot; The first quartile of the time taken for segment j in the box plot; Step 3044: Construct a subset M′ of vehicles on the road segment that does not include vehicles stopping at service areas. Where M is the original vehicle set, and M′ is the vehicle set after removing vehicles that have stopped at service areas; Step 3045: Further clean M′ to obtain vehicles affected by traffic anomalies, and then construct the segment abnormal vehicle subset I. j ′: in, The percentage of time taken for abnormal vehicles in a given section; This is the third quartile in the box plot; The interquartile interval in the box plot; This is the first quartile in the box plot; Step 3046, thus obtaining the effective vehicle subset M of the road segment: The effective subset of vehicles is the set of vehicles on the road segment used.

4. The method for measuring highway section mileage based on ETC big data according to claim 3, characterized in that: The abnormal traffic factors in step 3045 include traffic flow, road maintenance, and emergencies.

5. The method for measuring highway section mileage based on ETC big data according to claim 1, characterized in that: In step 4, the law of large numbers is used to construct a gantry section mileage generation model; Step 401: Based on the segment mileage obtained in Step 3, it can be seen that the mileage of each vehicle in different segments is independent of each other and follows the same distribution with mathematical expectation. in, μ represents the mileage traveled by the i-th vehicle in segment j. j Let be the mathematical expectation of the mileage of road segment j. Step 402, according to the law of large numbers, the sequence is... Converges in probability to μ j ,Right now set up With variance According to the central limit theorem, Δd j sum Standardized variable Y m for: Among them, Y m The distribution function F m (x) satisfies for any x: Among them, F m Φ(x) is the distribution function, and Φ(x) is the standard normal distribution function. Step 403, when m is not less than the minimum value, Δd j After appropriate standardization, the mean of the distribution converges to a normal distribution. mean Approximately follows the mean μ j The variance is If the value follows a normal distribution, then a gantry section mileage generation model MGM based on ETC big data is constructed: ΔD~N(μ,Γ Δd )。 6. The method for measuring highway section mileage based on ETC big data according to claim 5, characterized in that: The minimum value in step 403 is 30, that is, m≥30.

Citation Information

Patent Citations

  • Road network operation evaluation method based on vehicle travel data

    CN102819955A

  • Congestion level creation method and congestion level creation device

    JP2008020948A