Dynamic weighing axle load data quality control method based on probability density function
The dynamic weighing axle load data quality control method based on probability density function solves the data quality problem of dynamic weighing equipment under environmental changes, realizes real-time monitoring and evaluation, ensures equipment stability and data accuracy, and supports road design and maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG EXPRESSWAY GRP CO LTD INNOVATION RES INST
- Filing Date
- 2025-07-21
- Publication Date
- 2026-07-21
Smart Images

Figure CN120974223B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent transportation infrastructure monitoring technology, and more specifically, to a method for quality control of dynamic weighing axle load data based on probability density functions. Background Technology
[0002] In road engineering, axle load spectrum is an important basis for evaluating pavement design and service life. In order to monitor pavement performance over a long period of time, a large number of sensors, such as Weigh-In-Motion (WIM) devices, have been deployed on roads to collect vehicle axle load data in real time. In the long-term data monitoring, the equipment was in good working condition in the early stage and the data quality was high, which provided a reliable foundation for the establishment of axle load spectrum.
[0003] However, as the equipment operates for longer periods, the quality of dynamic weighing data gradually declines due to factors such as road surface unevenness, vehicle speed changes, and temperature drift, resulting in abnormal fluctuations and systematic errors. In order to obtain high-quality data over a long period and ensure the accuracy of axle load spectrum and the stability of equipment operation, it is necessary to effectively control and evaluate the quality of the collected data.
[0004] In existing technologies, data quality control methods often employ fixed threshold methods, which tend to have the following drawbacks: First, the threshold setting relies on empirical values, meaning it cannot adapt to changes in different road conditions and equipment states; second, it lacks sensitivity to occasional outliers, which can easily lead to misjudgments or omissions, affecting data quality; and finally, it lacks the ability to dynamically model the data distribution characteristics, making it difficult to meet the accuracy requirements for long-term data quality control and axle load spectrum establishment.
[0005] Therefore, there is an urgent need for a method that can dynamically adapt to environmental changes, detect outliers in real time, and automatically and comprehensively evaluate data quality to ensure understanding of equipment operating status and obtain high-quality data to support long-term pavement performance evaluation.
[0006] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention
[0007] To address the problems in related technologies, this invention proposes a dynamic weighing axle load data quality control method based on probability density function, in order to overcome the aforementioned technical problems existing in the existing related technologies.
[0008] Therefore, the specific technical solution adopted by the present invention is as follows:
[0009] A dynamic weighing axle load data quality control method based on probability density function, comprising the following steps:
[0010] S1. Based on the verification technology, extract the initial benchmark data from the historical monitoring data, and establish the benchmark probability density function of the axis type according to the initial benchmark data. Use the benchmark probability density function to establish the benchmark periodic characteristic parameters.
[0011] S2. Acquire axle load data for the monitoring period, perform quality control processing on the axle load data, classify the axle type of the output complete axle load data, and establish the probability density function of the test data based on the axle type classification results.
[0012] S3. Calculate the data distribution characteristics of the complete axle load data using the probability density function of the data to be tested, and determine the similarity between the initial reference data and the complete axle load data based on the data distribution characteristics and the reference period characteristic parameters;
[0013] S4. Compare the similarity results with the selected threshold requirements, and update and incorporate the axle load data of the monitoring period into the historical monitoring data based on the comparison results. Perform quality control on the axle load data of subsequent monitoring periods to generate dynamic closed-loop optimization.
[0014] Furthermore, based on verification techniques, initial baseline data is extracted from historical monitoring data, and a baseline probability density function for the axis type is established based on the initial baseline data. The baseline periodic characteristic parameters are then established using the baseline probability density function, including:
[0015] S11. Based on the verification technology, the previously acquired historical monitoring data is verified, the verified historical monitoring data is analyzed for time series characteristics, and the optimal analysis period is determined by combining the adaptive sliding window algorithm.
[0016] S12. Based on the optimal analysis cycle, extract the historical monitoring data after verification to obtain the axle load data of the previous cycle;
[0017] S13. Based on the preceding periodic axle load data, construct the benchmark probability density function of the axle type using the kernel density estimation method, and establish the benchmark periodic characteristic parameters through the benchmark probability density function.
[0018] The baseline periodic characteristic parameters include: dynamic axle load threshold, distribution pattern, central tendency, dispersion degree, and axle load heavy load ratio.
[0019] Furthermore, the dynamic axle load threshold is determined by the quantile method of the baseline probability density function, thus obtaining the data range defined by the upper and lower quantiles;
[0020] The formula for calculating the axle load heavy load ratio is as follows:
[0021] ;
[0022] In the formula, p_heavy1 represents the proportion of axle load under heavy load; a represents the lower limit of the heavy load range; and f1(x) represents the baseline probability density function.
[0023] Furthermore, axle load data for the monitoring period is acquired, quality control processing is performed on the axle load data, and axle type classification is performed on the output complete axle load data. Based on the axle type classification results, a probability density function for the test data is established, including:
[0024] S21. The axle load data of the pre-acquired monitoring period is subjected to integrity quality control using the first-level integrity quality verification mechanism. Based on the integrity quality control results, the axle load data of the monitoring period is subjected to integrity screening and abnormal alarms to obtain primary axle load data and axle load data with integrity abnormalities.
[0025] S22. Use a two-level consistency quality check mechanism to perform consistency quality control on the primary axle load data, and perform consistency screening and anomaly alarm on the primary axle load data based on the consistency quality control results to obtain intermediate axle load data and consistency anomaly axle load data.
[0026] S23. Combining dynamic axle load thresholds, a three-level rationality quality verification mechanism is used to perform rationality quality control on intermediate axle load data. Based on the rationality quality control results, rationality screening and abnormal alarms are performed on intermediate axle load data to obtain complete axle load data and rationality abnormal axle load data.
[0027] S24. Generate axle load data quality control log based on the quality control results, and perform axle type classification on the complete axle load data to obtain the axle type classification results. Establish the probability density function of the test data based on the axle type classification results.
[0028] Furthermore, the quality control results also include:
[0029] The data volume of axle load data with integrity abnormalities, consistency abnormalities, and rationality abnormalities is obtained separately, and the ratio is calculated with the data volume of axle load data, primary axle load data, and intermediate axle load data of the corresponding monitoring period to obtain the abnormality rate of each level. The equipment operating status is evaluated based on the results of the abnormality rate of each level.
[0030] Furthermore, the first-level integrity quality verification mechanism includes: identifying and marking axle load data with missing values in monitoring periods through missing value checks.
[0031] Furthermore, the secondary consistency quality verification mechanism includes: based on a preset axle load-vehicle type-number of axle association rules, combined with logical judgment to identify and mark axle load data with inconsistent performance.
[0032] Furthermore, the three-level rationality quality verification mechanism includes: calculating the axle load of the intermediate axle load data, and combining it with the dynamic axle load threshold to determine whether the axle load of the intermediate axle load data is a low-probability event. If it is a low-probability event, the intermediate axle load data is marked as axle load data with abnormal rationality.
[0033] Furthermore, the data distribution characteristics of the complete axle load data are calculated using the probability density function of the test data, and the similarity between the initial reference data and the complete axle load data is determined based on the data distribution characteristics and the reference period characteristic parameters, including:
[0034] S31. Calculate the data distribution characteristics of complete axle load data using the probability density function of the data to be tested.
[0035] S32. Using axle load distribution drift detection, the data distribution characteristics of complete axle load data and the reference period characteristic parameters are compared and analyzed to obtain the axle load distribution drift detection results.
[0036] S33. Obtain the changing trend of axle load data distribution based on the axle load distribution drift detection results, and evaluate the pressure change of axle load on the road surface through the changing trend;
[0037] S34. Calculate the axle load distribution similarity based on the data distribution characteristics of the complete axle load data, and use the axle load distribution similarity results to determine the similarity between the initial reference data and the complete axle load data.
[0038] Furthermore, the formula for calculating the similarity of axis weight distribution is as follows:
[0039] ;
[0040] In the formula, S represents the axle load distribution similarity; f1(x) represents the data distribution characteristics of the axle type; f2(x) represents the data distribution characteristics of the complete axle load data; min(f1(x),f2(x)) represents the intersection probability of the initial reference data and the complete axle load data.
[0041] The beneficial effects of this invention are as follows:
[0042] 1. This invention, through a baseline probability density function and a test data probability density function, can monitor and evaluate the quality of axle load data over a time range in real time and periodically. It can automatically and quickly detect abnormal axle load data, clarify the equipment operating status, ensure that complete axle load data is obtained for long-term pavement performance observation, and establish an axle load spectrum that accurately reflects road traffic conditions, providing a reliable basis for pavement design and structural maintenance. It solves the problems of existing data quality control methods that rely on fixed thresholds, cannot adapt to environmental changes, have insufficient sensitivity to outliers, and lack dynamic modeling capabilities.
[0043] 2. This invention, through a quality verification mechanism, can analyze the historical monitoring data collected by the dynamic weighing equipment one by one, mark abnormal axle load data, and save the verified axle load data, so that the quality of axle load data is controllable; and by analyzing the axle load data of the monitoring period, such as the abnormality rate of each level and the data distribution characteristics of complete axle load data, it can understand the changing trend of equipment condition and axle load data distribution, and dynamically update the benchmark. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0045] Figure 1 This is a flowchart of a dynamic weighing axle load data quality control method based on a probability density function according to an embodiment of the present invention;
[0046] Figure 2 This is a schematic diagram of the dynamic axle load threshold obtained by the reference probability density function in the dynamic weighing axle load data quality control method based on the probability density function according to an embodiment of the present invention.
[0047] Figure 3 This is a schematic diagram of the multi-dimensional benchmark index obtained by the benchmark probability density function in the dynamic weighing axle load data quality control method based on the probability density function according to an embodiment of the present invention.
[0048] Figure 4 This is a schematic diagram of axle load distribution similarity in a dynamic weighing axle load data quality control method based on probability density function according to an embodiment of the present invention. Detailed Implementation
[0049] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.
[0050] According to an embodiment of the present invention, a method for quality control of dynamic weighing axle load data based on probability density function is provided.
[0051] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 As shown, the dynamic weighing axle load data quality control method based on probability density function according to an embodiment of the present invention includes the following steps:
[0052] S1. Based on the verification technology, extract the initial benchmark data from the historical monitoring data, and establish the benchmark probability density function of the axis type according to the initial benchmark data. Use the benchmark probability density function to establish the benchmark periodic characteristic parameters.
[0053] Specifically, initial baseline data is extracted from historical monitoring data based on verification techniques, and a baseline probability density function for the axis type is established based on the initial baseline data. The baseline periodic characteristic parameters are then established using the baseline probability density function, including:
[0054] S11. Based on the verification technology, the previously acquired historical monitoring data is verified, the verified historical monitoring data is analyzed for time series characteristics, and the optimal analysis period is determined by combining the adaptive sliding window algorithm.
[0055] S12. Based on the optimal analysis cycle, extract the historical monitoring data after verification to obtain the axle load data of the previous cycle;
[0056] S13. Based on the preceding periodic axle load data, construct the benchmark probability density function of the axle type using the kernel density estimation method, and establish the benchmark periodic characteristic parameters through the benchmark probability density function.
[0057] The baseline periodic characteristic parameters include: dynamic axle load threshold, distribution pattern, central tendency, dispersion degree, and axle load heavy load ratio.
[0058] Specifically, the dynamic axle weight threshold is determined by the quantile method of the baseline probability density function, thus obtaining the data range defined by the upper and lower quantiles;
[0059] Formula for calculating the proportion of axle load under heavy load:
[0060] ;
[0061] In the formula, p_heavy represents the proportion of axle load under heavy load; a represents the lower limit of the heavy load range; f(x) represents the baseline probability density function; and x represents the axle load data of the preceding period.
[0062] Formula for calculating dynamic axle load threshold:
[0063] ;
[0064] In the formula, Indicates the dynamic axle load threshold; Indicates quantiles; The cumulative distribution function represents the baseline probability density function.
[0065] Specifically, the data distribution characteristics of axis types also include the median axis weight and the mode axis weight.
[0066] Formula for calculating the median axle load:
[0067] ;
[0068] In the formula, Med represents the median of the baseline axle load.
[0069] Formula for calculating the mode of axle load:
[0070] ;
[0071] In the formula, mode represents the mode of the reference axle weight; It represents the first derivative of the baseline probability density function.
[0072] Specifically, based on the verification technology, initial benchmark data is extracted from historical monitoring data, and a benchmark probability density function for the axis type is established. The benchmark probability density function is used to establish benchmark period characteristic parameters, which involves intelligent configuration of the analysis period and benchmark modeling. Based on the time series characteristic analysis of historical monitoring data, in the system initialization stage, the axis load data of the previous period is extracted, and the kernel density estimation method is used to construct the benchmark probability density function for the axis type. Multi-dimensional statistical indicators (i.e., data distribution characteristics of the axis type) including dynamic thresholds and distribution patterns are established as quantitative benchmarks for data quality evaluation in this period.
[0073] Specifically, the historical monitoring data obtained is data collected through dynamic weighing equipment, including: date, time, vehicle speed, vehicle type, number of axles, and axle load.
[0074] Specifically, the optimal analysis period is to perform periodic analysis on historical monitoring data, and the time range can be set to week, month, or quarter.
[0075] Specifically, the baseline probability density function (PDF) is established according to axle type, which is divided into four categories: single axle single wheel, single axle dual wheel, dual axle, and triple axle, based on the asphalt pavement design specifications of a certain year. All axle load monitoring data in the preceding reliable period (i.e., axle load data of the preceding period) are extracted, and all axle load monitoring data in the preceding reliable period includes single axle, dual axle, and multi-axle vehicle types. The kernel density estimation method (KDE) is used to construct the baseline probability density function (PDF) for each axle type as the benchmark for the distribution of axle load of each type in the monitoring period, denoted as f(x).
[0076] The baseline probability density function (PDF) is based on the baseline axle load data with confirmed quality (i.e., the axle load data of the preceding period). The PDF is established using kernel density estimation (KDE) and serves as the benchmark for the load distribution of each type of axle in the monitoring period, denoted as f(x). The axle load characteristic information of the baseline axle load data is all included in the PDF.
[0077] Specifically, the low-probability event threshold (i.e., the dynamic axis weight threshold) is determined using the quantile method of the baseline probability density function. This involves defining a reasonable data range using upper and lower quantiles, where data outside this range are statistically significant. It also means less than The probability that the axle load is in the distribution is ; It also represents the integral of the baseline probability density function.
[0078] Specifically, the calculation of the heavy load ratio of the benchmark axle load uses the mathematical properties of the integral of the benchmark probability density function, and the heavy load ratio p_heavy of the benchmark axle load also represents the probability of being in the heavy load range; 'a' represents the lower limit of the heavy load range. Heavy load plays a key role in road damage. The lower limit of the heavy load range is divided according to the vehicle axle load overload identification standard as follows: 60kN for single axle single wheel, 100kN for single axle double wheel, 180kN for double axle and 220kN for triple axle.
[0079] S2. Acquire axle load data for the monitoring period, perform quality control processing on the axle load data, classify the axle type of the output complete axle load data, and establish the probability density function of the test data based on the axle type classification results.
[0080] Specifically, the process involves acquiring axle load data for the monitoring period, performing quality control processing on the axle load data, classifying the axle type of the output complete axle load data, and establishing the probability density function of the test data based on the axle type classification results, including:
[0081] S21. The axle load data of the pre-acquired monitoring period is subjected to integrity quality control using the first-level integrity quality verification mechanism. Based on the integrity quality control results, the axle load data of the monitoring period is subjected to integrity screening and abnormal alarms to obtain primary axle load data and axle load data with integrity abnormalities.
[0082] S22. Use a two-level consistency quality check mechanism to perform consistency quality control on the primary axle load data, and perform consistency screening and anomaly alarm on the primary axle load data based on the consistency quality control results to obtain intermediate axle load data and consistency anomaly axle load data.
[0083] Specifically, a quality verification mechanism is adopted to analyze the axle load data collected by the dynamic weighing equipment during the monitoring period, mark abnormal data, and save the verified data, so that the quality of axle load data is controllable.
[0084] Specifically, line-by-line analysis is a real-time data analysis method, which involves conducting real-time analysis of vehicle information acquired by dynamic weighing equipment.
[0085] Specifically, the first-level integrity quality verification mechanism includes: identifying and marking axle load data with missing values in monitoring periods through missing value checks.
[0086] Specifically, the missing value check automatically marks any missing data records as abnormal. This means that all elements of the axle load data in the monitoring period must be recorded; otherwise, it is abnormal. Elements include date, time, vehicle speed, vehicle type, number of axles, and axle weight. It mainly analyzes the completeness of the axle load data in the monitoring period. By identifying whether there are null values, it checks whether each element of each data record has recorded information. If there are null values, the data is marked as abnormal and will not participate in the secondary consistency quality check.
[0087] Specifically, the secondary consistency quality verification mechanism includes: based on a preset axle load-vehicle type-number of axle association rules, combined with logical judgment to identify and mark axle load data with inconsistent performance.
[0088] Specifically, the Level 2 consistency quality check requires analyzing the consistency of the primary axle load data. Based on a preset axle weight-vehicle type-number of axle association rule base, it performs logical judgment to identify conflicting data (i.e., Level 2 abnormal axle load data). This preset rule base is derived from vehicle type classification rules. For example, if the recorded vehicle type is "Type 157", it should theoretically have 6 axles. If the axle load data only contains the weight of 5 axles, it is inferred from the preset rule base that the data information is inconsistent. Through logical judgment, this data is marked as inconsistent axle load data and will not participate in the Level 3 rationality quality check.
[0089] Specifically, the consistency of primary axle load data refers to the reasonable coordination among the relevant elements of the primary axle load data. For example, the theoretical number of axles corresponding to the vehicle model and the actual number of axles monitored should be consistent; otherwise, it is abnormal.
[0090] S23. Combining dynamic axle load thresholds, a three-level rationality quality verification mechanism is used to perform rationality quality control on intermediate axle load data. Based on the rationality quality control results, rationality screening and abnormal alarms are performed on intermediate axle load data to obtain complete axle load data and rationality abnormal axle load data.
[0091] S24. Generate axle load data quality control log based on the quality control results, and perform axle type classification on the complete axle load data to obtain the axle type classification results. Establish the probability density function of the test data based on the axle type classification results.
[0092] Specifically, the three-level rationality quality verification mechanism includes: calculating the axle load of the intermediate axle load data, and combining the dynamic axle load threshold to determine whether the axle load of the intermediate axle load data is a low-probability event. If it is a low-probability event, the intermediate axle load data is marked as axle load data with abnormal rationality.
[0093] Specifically, the intermediate axle load data is subjected to reasonableness quality control by combining the dynamic axle load threshold. That is, the axle load of the intermediate axle load data is compared with the dynamic axle load threshold. When the axle load is less than the dynamic axle load threshold, it is marked as "abnormal to be verified", that is, axle load data with reasonableness abnormality. At this time, the intermediate axle load data is a low probability event, that is, whether the data is outside the dynamic axle load threshold range of the benchmark data is questionable and needs to be further manually verified.
[0094] Among them, data outside the dynamic axle weight threshold range are statistically significant. Statistical significance means that the probability of an event occurring outside the range is very small, but it still occurs. Such low-probability events may be outliers, and an anomaly prompt needs to be sent to the front end.
[0095] The possible abnormal situations are represented as follows:
[0096] ;
[0097] In the formula, x t The axle load represents the axle load data of a certain type of axle under real-time monitoring; This indicates the lower bound of the axis weight of this type of axis in the baseline probability density function; It is the weight of the upper bound of the probability density function for this type of axis.
[0098] Specifically, the quality control results also include:
[0099] The data volume of axle load data with integrity abnormalities, consistency abnormalities, and rationality abnormalities is obtained separately, and the ratio is calculated with the data volume of axle load data, primary axle load data, and intermediate axle load data of the corresponding monitoring period to obtain the abnormality rate of each level. The equipment operating status is evaluated based on the results of the abnormality rate of each level.
[0100] S3. Calculate the data distribution characteristics of the complete axle load data using the probability density function of the test data, and determine the similarity between the initial reference data and the complete axle load data based on the data distribution characteristics and the reference period characteristic parameters.
[0101] Specifically, after data quality control processing, the axle load data of the monitoring cycle is analyzed, including the statistical abnormality rate of each level of data within the cycle, the calculation of data distribution characteristics of the monitoring cycle data and axle load distribution drift detection, etc., to clarify the equipment operating status and data development trend; among them, after the data collection of the cycle is completed and the potentially abnormal data is investigated, a comprehensive quality analysis and feature analysis of the data within the cycle is automatically performed to understand the equipment operating status and data changes, that is, to perform timed analysis of complete axle load data.
[0102] Specifically, the data distribution characteristics of the complete axle load data are calculated using the probability density function of the test data, and the similarity between the initial reference data and the complete axle load data is determined based on the data distribution characteristics and the reference period characteristic parameters, including:
[0103] S31. Calculate the data distribution characteristics of complete axle load data using the probability density function of the test data.
[0104] Specifically, the formulas for calculating the abnormality rate at each level are as follows:
[0105] ;
[0106] In the formula, R represents the anomaly rate; Indicates abnormal axle load data volume; This indicates the total amount of axle load data under each level of quality verification.
[0107] Specifically, the abnormal axle load data volume refers to the data volume of axle load data with abnormal integrity, abnormal consistency, or abnormal rationality; the total axle load data volume under each level of quality verification refers to the data volume of axle load data during the monitoring period, primary axle load data, or intermediate axle load data.
[0108] Specifically, the anomaly rates are statistically graded by calculating the anomaly rates at each level. The anomaly rates at each level include the integrity verification anomaly rate, the consistency verification anomaly rate, and the rationality verification anomaly rate. All anomaly rates need to be less than 2%. The integrity verification anomaly rate is obtained by dividing the amount of integrity-abnormal axle load data by the amount of axle load data in the monitoring period. The consistency verification anomaly rate is obtained by dividing the amount of consistency-abnormal axle load data by the amount of primary axle load data. The rationality verification anomaly rate is obtained by dividing the amount of rationality-abnormal axle load data by the amount of intermediate axle load data.
[0109] Specifically, the operating status of equipment is assessed by using anomaly rates at various levels. For example, the integrity verification anomaly rate can directly reflect the operating status of each component of the weighing equipment. An increase in the integrity verification anomaly rate indicates that there are component failures or performance degradation in the equipment.
[0110] S32. Using axle load distribution drift detection, the data distribution characteristics of complete axle load data and the reference period characteristic parameters are compared and analyzed to obtain the axle load distribution drift detection results.
[0111] S33. Obtain the changing trend of axle load data distribution based on the axle load distribution drift detection results, and evaluate the pressure change of axle load on the road surface through the changing trend.
[0112] S34. Calculate the axle load distribution similarity based on the data distribution characteristics of the complete axle load data, and use the axle load distribution similarity results to determine the similarity between the initial reference data and the complete axle load data.
[0113] Specifically, the formula for calculating the similarity of axis weight distribution is as follows:
[0114] ;
[0115] In the formula, S represents the axle load distribution similarity; f1(x) represents the data distribution characteristics of the axle type; f2(x) represents the data distribution characteristics of the complete axle load data; min(f1(x),f2(x)) represents the intersection probability of the initial reference data and the complete axle load data.
[0116] Specifically, axle load distribution similarity is represented by the common area of the two probability density functions within the same interval, i.e., the overlap of the probability distributions, which is the axle load distribution similarity calculated by the data distribution characteristics of the complete axle load data and the data distribution characteristics of the axle shape. The similarity between the initial baseline data and the complete axle load data is calculated based on the two probability density functions, and min(f1(x), f2(x)) also represents the minimum value of the two PDFs at each point. Here, axle load distribution similarity is an indicator set to clarify the rationality of the data distribution, and can be intuitively understood as the overlap of the probability distributions in a graphical representation. Therefore, similarity is expressed using two... The common area of the probability density functions of the same axle type in different periods is used to represent this; the similarity of the data distributions of the two periods is calculated based on the two probability density functions; regarding the criteria for evaluating good similarity, based on the experience of MEPDG (as shown in Table 1); a 95% confidence level for the error obtained by comparing 30 days of axle load data with one year of data within a 2% range means that the 30-day data can reflect the axle load distribution of the whole year well; according to this experience, if the overlap between the monthly data of each period and the annual data can reach 0.98 × 0.95 = 0.931, it indicates that the data has good similarity distribution and substitutability. And if the overlap of the probability density functions between two monthly periods reaches 0.931 × 0.931 = 0.867, it is likely that both can reflect the axle load distribution of the year well. The threshold for other periods can also be determined according to this idea.
[0117] Table 1 Minimum sample size (days / year) required to estimate axle load distribution
[0118] S4. Compare the similarity results with the selected threshold requirements, and update and incorporate the axle load data of the monitoring period into the historical monitoring data based on the comparison results. Perform quality control on the axle load data of subsequent monitoring periods to generate dynamic closed-loop optimization.
[0119] Specifically, the similarity results are compared with the selected threshold requirements to determine whether the similarity meets the baseline update conditions. If the anomaly rate at each level is less than or equal to 2% and the axle weight distribution similarity meets certain conditions, the baseline update program is initiated to update the baseline probability density function. The updated baseline probability density function synchronously optimizes the dynamic axle weight threshold for anomaly judgment in the next cycle, forming a closed-loop optimization of quality control parameters.
[0120] In summary, this invention establishes a baseline probability density function based on kernel density estimation to monitor and periodically analyze axle load data collected by dynamic weighing equipment in real time, ensuring controllable axle load data quality and normal equipment operation. Specifically, it includes: axle load data screening and preprocessing, axle load data quality control, and a dynamic axle load data update mechanism. This invention can dynamically adapt to environmental changes, ensuring long-term data quality and equipment operating status, providing a reliable basis for pavement design and structural maintenance. Furthermore, this invention is particularly suitable for real-time data quality control of Weigh-In-Motion (WIM) equipment in highway traffic, aiming to provide accurate data support for the establishment of axle load spectra in road engineering.
[0121] To facilitate understanding of the above technical solutions of the present invention, the working principle or operation method of the present invention in actual process will be described in detail below.
[0122] In practical applications, taking a long-term performance observation point in a certain region as an example, the dynamic weighing equipment can monitor axle load information in real time. The information collected includes date, time, vehicle speed, vehicle type, number of axles, and axle weight.
[0123] Step 1: Determine the analysis period. In this example, the period is monthly. The historical dynamic weighing axle load monitoring data of the previous month is used as the initial reference data. This initial reference data is manually verified. Based on the reference data (i.e., the axle load data of the previous period), the reference probability density function of each axle type is established as the reference for the example month.
[0124] The choice of dynamic axle weight threshold quantiles depends on the situation. For stricter monitoring, a wider quantile can be selected. In this example, 0.01 is used as the lower threshold quantile and 0.99 as the upper threshold quantile. A schematic diagram of calculating the dynamic axle weight threshold using the baseline probability density function for establishing the axle type is shown (e.g.). Figure 2 (As shown).
[0125] The median and mode of the axis weight distribution are calculated using the baseline probability density function used to establish the axis type (e.g., ...). Figure 3 (As shown).
[0126] The probability density function of the baseline axle type is used to calculate the probability of heavy load distribution areas (i.e., the proportion of axle load heavy load) for different axle types (e.g., ...). Figure 3As shown in the figure, the lower limit of the heavy load range is selected according to the vehicle axle load overload identification standard: 60kN for single axle single wheel, 100kN for single axle double wheel, 180kN for double axle and 220kN for triple axle.
[0127] Based on the baseline probability density function of the axis type, according to Figure 2 , Figure 3 The dynamic axle load threshold, median, and mode of axle load obtained from the schematic diagram are shown in Table 2, which classifies the axle load by axle type.
[0128] Table 2. Dynamic axle load thresholds and eigenvalues calculated based on the baseline probability density function of the axle type.
[0129] Step two: Implement a quality verification mechanism during the monitoring period. This mechanism includes a first-level integrity quality verification mechanism, a second-level consistency quality verification mechanism, and a third-level rationality quality verification mechanism. All abnormal axle load data will trigger tiered alarms and generate quality control logs. The specific steps are as follows:
[0130] First, a first-level integrity quality check is performed, requiring all elements of each data entry to be recorded without any missing or null values. Real-time analysis is then performed in chronological order to evaluate the integrity of each data entry. Data that does not meet the integrity requirements is recorded separately in the integrity check database for backup, preventing contamination of subsequent data. Tables 3.1 and 3.2 list some data that do not meet the integrity requirements within the analysis period, such as missing elements like time, vehicle speed, and vehicle type. N represents data that does not meet the integrity requirements.
[0131] Table 3.1 Demonstration of Level 1 Integrity Verification Results
[0132] Table 3.2 Demonstration of Level 1 Integrity Verification Results
[0133] Secondly, a second-level consistency quality check mechanism is implemented for data that has passed the first-level integrity quality check without any missing data. This involves combining the vehicle type-axle number-axle load column for logical judgment, identifying conflicting data, recording it separately in the consistency check data frame for backup, and preventing it from contaminating subsequent data.
[0134] Tables 4.1 and 4.2 list some data that do not meet the consistency requirement within the analysis period. For example, the data recorded at 13:03:44 on a certain day of the month shows that it is a Type 157 vehicle.
[0135] According to the "Asphalt Pavement Design Specification", the vehicle composition theoretically has 6 axles, but the recorded number of axles is actually only 5, which causes information conflict. N represents the data that does not meet the consistency requirements.
[0136] Table 4.1 Demonstration of Level 2 Consistency Check Results
[0137] Table 4.2 Demonstration of Level 2 Consistency Check Results
[0138] Finally, a third-level rationality quality check is performed on the data that has passed the first two levels of verification. That is, according to the single axle load of the vehicle model combination, the axle load data is divided into four categories of axles: single axle single wheel, single axle double wheel, double axle and triple axle. The rationality of the axle load data is evaluated based on the dynamic axle load threshold of the baseline probability density function of each category of axles in the baseline data.
[0139] Tables 5.1 and 5.2 show the process of axle load combination and the threshold judgment results. N indicates that the threshold requirements are not met and further manual verification is needed. As can be seen from Tables 5.1 and 5.2, this information has no missing values, meeting the integrity check. Theoretically, the 115-type truck has 4 axles matching the detected axle number, meeting the consistency check. Therefore, it proceeds to the rationality check. The axle load data is combined into two "Type 1 axles" (single axle, single wheel) and one "Type 5 axle" (double axle) according to "115". The thresholds extracted from the baseline probability density function in Table 2 show that the normal range for a single axle, single wheel is 3.92kN to 76.05kN, and the normal range for a double axle is 20.58kN to 254.21kN. Therefore, it is determined that neither of the two single axle, single wheel axles is within the threshold range, while the double axle is within the threshold range.
[0140] Table 5.1 Demonstration of the effect of axle load combination and dynamic axle load threshold judgment
[0141] Table 5.2 Demonstration of the effect of axle load combination and dynamic axle load threshold judgment
[0142] The combined axle load data is categorized by the combined axle type for rationality assessment. Information that does not meet the threshold is recorded and awaits manual verification. Weighing equipment typically records license plate numbers, which staff can use to verify the data against other nearby weighing equipment. Tables 6.1 and 6.2 show the recorded data exceeding the threshold range and the verification process. If the rationality verification result is Y, the data is saved to the final database; otherwise, it is recorded in the rationality verification exception log.
[0143] This dynamic axle load threshold combined with manual verification allows for flexible data filtering without damaging the data. For example, according to the vehicle axle load overload determination standard, a single axle or single wheel load of 60kN or more is considered overloaded. The static threshold may not exceed 120kN. Traditional static thresholds combined with direct rejection methods would directly reject data like 162.288kN, but such data may still exist and have a greater impact on the road surface. Direct rejection would affect road performance analysis and hinder more accurate design and maintenance.
[0144] Table 6.1 Display of Axle Load Data Threshold Judgment and Preparation for Reasonableness Verification
[0145] Table 6.2 Display of Axle Load Data Threshold Judgment and Preparation for Reasonableness Verification
[0146] Step 3: After data quality control is completed, data without anomalies will be saved to a data frame. Subsequent automatic analysis of the data within the period is possible, including anomaly rate, distribution drift detection, and clarification of equipment operating status and data trends. This includes the following steps:
[0147] First, the anomaly rate of the data within the period is calculated. This anomaly rate includes the anomaly rate of each level in the three-level quality control. In this example, the Level 1 integrity quality check mechanism collected 130,435 data entries, of which 216 data entries did not meet the integrity requirements, resulting in an anomaly rate of 0.16%. The Level 2 consistency quality check mechanism received 130,219 input data entries, of which 107 data entries did not meet the consistency requirements, resulting in an anomaly rate of 0.08%. The Level 3 rationality quality check mechanism received 231,771 single-axis single-wheel data entries, 17,252 single-axis double-wheel data entries, 16,963 double-axis data entries, and 13,384 triple-axis data entries. After dynamic axle weight threshold judgment and manual verification, the abnormal data were 63 single-axis single-wheel data entries, 17 single-axis double-wheel data entries, 12 double-axis data entries, and 24 triple-axis data entries, with anomaly rates of 0.02%, 0.09%, 0.7%, and 0.18%, respectively.
[0148] Secondly, kernel density estimation was used to establish the probability density function of the monitoring period for the complete axle load data, and the distribution of axle load data in the monitoring period was calculated, including the median axle load (mode), the mode of axle load (Med), and the proportion of heavy load (p_heavy) of axle load in the monitoring period. Heavy load was defined according to the vehicle axle load overload identification standards: 60kN for a single axle with a single wheel, 100kN for a single axle with two wheels, 180kN for a double axle, and 220kN for a triple axle. The results are shown in Table 7. Comparative analysis with the baseline data in Table 2 reveals that the high-frequency axle loads of most axles increased during the monitoring period, but the median axle load of most axles decreased, and the proportion of heavy axles also decreased. This indicates that the axle load in the monitoring period decreased compared to the baseline period, suggesting a possible reduction in road surface load pressure.
[0149] Table 7. Data distribution characteristics of complete axle load data within the monitoring period. Finally, according to Figure 4 As shown in Table 8, the similarity between the initial baseline data and the complete axle load data distribution, calculated using the similarity calculation formula, yields the following results. Table 8 shows that the overlap between the probability density functions of various axles during the monitoring period and the baseline period is at least 0.8925, meaning that 89.25% of the area overlaps.
[0150] Regarding the criteria for evaluating good similarity, based on MEPDG's experience (as shown in Table 1), a 95% confidence level is achieved when the error obtained by comparing 30 days of axle load data with one year's data is within a 2% range. This means that 30 days of data can reflect the axle load distribution throughout the year quite well. Based on this experience, if the overlap between the monthly data and the annual data reaches 0.98 × 0.95 = 0.931, it indicates good similarity in distribution and substitutability. Furthermore, if the overlap of the probability density functions between two monthly periods reaches 0.931 × 0.931 = 0.867, it likely reflects the axle load distribution throughout the year well. Therefore, this example uses 0.867 as the similarity threshold.
[0151] Table 8. Similarity Judgment Results
[0152] Step 4, Updateability Analysis: Based on the findings from Step 3, all anomaly rates at each level are less than 2%, and the similarity meets the threshold requirements of the selected preset rules, indicating sufficient correlation. Therefore, the benchmark update program can be initiated. The probability density function of the monitoring period of normal data in the monitoring period will be used as the benchmark for the next period. The benchmark data will be iteratively updated, and the quality of the new data will be controlled to form a closed-loop optimization of the quality control parameters.
[0153] In summary, by utilizing the above-mentioned technical solution of this invention, the present invention, through the benchmark probability density function and the probability density function of the data to be tested, can monitor and evaluate the quality of axle load data over a time range in real time and periodically. It can automatically and quickly detect abnormal axle load data, clarify the equipment operating status, ensure that complete axle load data is obtained for long-term pavement performance observation, and establish an axle load spectrum that can accurately reflect road traffic conditions, providing a reliable basis for pavement design and structural maintenance. It solves the problems of existing data quality control methods that rely on fixed thresholds, cannot adapt to environmental changes, have insufficient sensitivity to outliers, and lack dynamic modeling capabilities. Through a quality verification mechanism, this invention can analyze the historical monitoring data collected by dynamic weighing equipment line by line, mark abnormal axle load data, and save verified axle load data, making the quality of axle load data controllable. Furthermore, by analyzing the axle load data during the monitoring period, such as the abnormality rate at each level and the data distribution characteristics of complete axle load data, it can understand the changing trends of equipment status and axle load data distribution, and dynamically update the benchmark.
[0154] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for quality control of dynamic weighing axle load data based on probability density function, characterized in that, The method includes: S1. Based on the verification technology, extract the initial benchmark data from the historical monitoring data, and establish the benchmark probability density function of the axis type according to the initial benchmark data. Use the benchmark probability density function to establish the benchmark periodic characteristic parameters. S2. Acquire axle load data for the monitoring period, perform quality control processing on the axle load data, classify the axle type of the output complete axle load data, and establish a probability density function for the data to be tested based on the axle type classification results; S2 includes: S21. The axle load data of the pre-acquired monitoring period is subjected to integrity quality control using the first-level integrity quality verification mechanism. Based on the integrity quality control results, the axle load data of the monitoring period is subjected to integrity screening and abnormal alarms to obtain primary axle load data and axle load data with integrity abnormalities. S22. Use a two-level consistency quality check mechanism to perform consistency quality control on the primary axle load data, and perform consistency screening and anomaly alarm on the primary axle load data based on the consistency quality control results to obtain intermediate axle load data and consistency anomaly axle load data. S23. Combining dynamic axle load thresholds, a three-level rationality quality verification mechanism is used to perform rationality quality control on intermediate axle load data. Based on the rationality quality control results, rationality screening and abnormal alarms are performed on intermediate axle load data to obtain complete axle load data and rationality abnormal axle load data. S24. Generate axle load data quality control log based on the quality control results, and perform axle type classification on the complete axle load data to obtain the axle type classification results. Establish the probability density function of the data to be tested based on the axle type classification results. S3. Calculate the data distribution characteristics of the complete axle load data using the probability density function of the data to be tested, and determine the similarity between the initial reference data and the complete axle load data based on the data distribution characteristics and the reference period characteristic parameters; S3 includes: S31. Calculate the data distribution characteristics of complete axle load data using the probability density function of the data to be tested. S32. Using axle load distribution drift detection, the data distribution characteristics of complete axle load data and the reference period characteristic parameters are compared and analyzed to obtain the axle load distribution drift detection results. S33. Obtain the changing trend of axle load data distribution based on the axle load distribution drift detection results, and evaluate the pressure change of axle load on the road surface through the changing trend; S34. Calculate the axle load distribution similarity based on the data distribution characteristics of the complete axle load data, and use the axle load distribution similarity results to determine the similarity between the initial reference data and the complete axle load data; S4. Compare the similarity results with the selected threshold requirements, and update and incorporate the axle load data of the monitoring period into the historical monitoring data based on the comparison results. Perform quality control on the axle load data of subsequent monitoring periods to generate dynamic closed-loop optimization.
2. The dynamic weighing axle load data quality control method based on probability density function according to claim 1, characterized in that, The method involves extracting initial baseline data from historical monitoring data using verification techniques, establishing a baseline probability density function for the axis type based on the initial baseline data, and using the baseline probability density function to establish baseline periodic characteristic parameters, including: S11. Based on the verification technology, the previously acquired historical monitoring data is verified, the verified historical monitoring data is analyzed for time series characteristics, and the optimal analysis period is determined by combining the adaptive sliding window algorithm. S12. Based on the optimal analysis cycle, extract the historical monitoring data after verification to obtain the axle load data of the previous cycle; S13. Based on the preceding periodic axle load data, construct the benchmark probability density function of the axle type using the kernel density estimation method, and establish the benchmark periodic characteristic parameters through the benchmark probability density function. The baseline periodic characteristic parameters include: dynamic axle load threshold, distribution pattern, central tendency, dispersion degree, and axle load heavy load ratio.
3. The dynamic weighing axle load data quality control method based on probability density function according to claim 2, characterized in that, The dynamic axle load threshold is determined by the quantile method of the baseline probability density function, and the data range defined by the upper and lower quantiles is obtained. The formula for calculating the axle load heavy load ratio is as follows: ; In the formula, p_heavy represents the proportion of axle load under heavy load; a represents the lower limit of the heavy load range; f(x) represents the baseline probability density function; and x represents the axle load data of the preceding period.
4. The dynamic weighing axle load data quality control method based on probability density function according to claim 1, characterized in that, The quality control results also include: The data volume of axle load data with integrity abnormalities, consistency abnormalities, and rationality abnormalities is obtained separately, and the ratio is calculated with the data volume of axle load data, primary axle load data, and intermediate axle load data of the corresponding monitoring period to obtain the abnormality rate of each level. The equipment operating status is evaluated based on the results of the abnormality rate of each level.
5. The dynamic weighing axle load data quality control method based on probability density function according to claim 1, characterized in that, The first-level integrity quality verification mechanism includes: identifying and marking axle load data with missing values in monitoring periods through missing value checks.
6. The dynamic weighing axle load data quality control method based on probability density function according to claim 1, characterized in that, The secondary consistency quality verification mechanism includes: based on a preset axle load-vehicle type-number of axle association rules, combined with logical judgment to identify and mark axle load data with inconsistent performance.
7. The dynamic weighing axle load data quality control method based on probability density function according to claim 1, characterized in that, The three-level rationality quality verification mechanism includes: calculating the axle load of the intermediate axle load data, and combining the dynamic axle load threshold to determine whether the axle load of the intermediate axle load data is a low-probability event. If it is a low-probability event, the intermediate axle load data is marked as axle load data with abnormal rationality.
8. The dynamic weighing axle load data quality control method based on probability density function according to claim 1, characterized in that, The formula for calculating the similarity of axis redistribution is as follows: ; In the formula, S represents the axle load distribution similarity; f1(x) represents the data distribution characteristics of the axle type; f2(x) represents the data distribution characteristics of the complete axle load data; min(f1(x),f2(x)) represents the intersection probability of the initial reference data and the complete axle load data.