Traffic big data cleaning system based on cloud computing

Through the cloud-based traffic big data cleaning system, the multi-level dynamic cleaning and quality evaluation module is adopted to solve the problems of single cleaning strategies and insufficient quality evaluation in the existing technology, and the intelligent adjustment and accurate evaluation of data quality are realized, and the cleaning effect is improved.

CN120492440APending Publication Date: 2025-08-15HUIZHOU SENYUAN INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510590534.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the existing technology, the transportation big data cleaning strategy is single, and it is unable to dynamically adapt to data changes. It lacks quality evaluation links, so it is impossible to verify whether the cleaning effect meets actual needs.

Method used

A cloud computing-based traffic big data cleaning system is designed, including data acquisition and distributed storage, multi-level dynamic cleaning and quality evaluation modules. Data clustering and abnormal detection are performed through DBSCAN, OPTICS and Xbar-S control chart algorithms, combined with particle swarm optimization LSTM model and support vector regression for cleaning, a quality template library is built for data quality evaluation, and a three-level cleaning mode is dynamically selected.

Benefits of technology

It realizes intelligent adjustment of data cleaning, improves the accuracy of data quality evaluation and the adaptability of cleaning strategies, reduces the risk of misjudgment, and forms a closed-loop feedback for data quality improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492440A_ABST
    Figure CN120492440A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic big data cleaning system based on cloud computing. The traffic big data cleaning system comprises a data acquisition and distributed storage module, a multi-stage dynamic cleaning module and a quality evaluation module, relates to the technical field of data processing, a data missing rate and an abnormal data proportion are calculated through a statistical analysis method, a three-level cleaning mode is dynamically selected, intelligent adjustment of the cleaning strength is achieved, and through setting in a quality evaluation module and calculation of space-time similarity between cleaning data and historical quality templates in a quality template library, the cleaning strength is evaluated. The cosine similarity and the attenuation factor based on the Euclidean distance are combined, the direction and distance information of the data are comprehensively considered, the similarity between the cleaned data and the historical quality template is evaluated from multiple dimensions, the one-sidedness of single index evaluation is avoided, cleaning strategy upgrading is supported, closed-loop feedback of data quality improvement is formed, and the quality of the data is improved. And the misjudgment risk of the cleaning strategy decision-making unit is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a traffic big data cleaning system based on cloud computing. Background Art

[0002] Data cleaning refers to the process of reviewing and verifying raw data to remove erroneous, duplicate, incomplete or inconsistent data in order to improve data quality and make it more suitable for subsequent analysis, mining and decision-making tasks.

[0003] Publication number CN106202335B discloses a traffic big data cleaning method based on a cloud computing framework. First, the entire data source is scanned. If there is missing data, it is filled based on the adjacent quadratic mean of the dimension where the data of the same road section is located; then, data with similar data change patterns are clustered into one category to obtain the cluster center of the road section data; finally, new data is matched with the cluster center number with the smallest distance to update or eliminate abnormal data.

[0004] However, the above application still has the following problems: the above application cleaning strategy is single, and only uses fixed cluster center matching, which cannot dynamically adapt to data changes, and the cleaned data lacks a quality assessment link, and it is impossible to verify whether the cleaning effect meets actual needs. Summary of the Invention

[0005] In order to solve the technical problems existing in the background technology, the present invention proposes a traffic big data cleaning system based on cloud computing.

[0006] The present invention proposes a cloud computing-based traffic big data cleaning system, comprising:

[0007] Data collection and distributed storage module: used to collect numerical data from multi-source traffic data in real time. The numerical data in multi-source traffic data includes location information data, speed data, traffic flow data and operation trajectory data of transportation vehicles;

[0008] By matching with the map database, the corresponding road section number information is added to each multi-source traffic data, and the multi-source traffic data is partitioned by road section number and stored in the distributed file system of the cloud platform;

[0009] Multi-level dynamic cleaning module: used to clean numerical data in multi-source traffic data;

[0010] Quality assessment module: used to perform quality assessment on data cleaned by the multi-level dynamic cleaning module.

[0011] Preferably, the multi-stage dynamic cleaning module includes:

[0012] Missing data processing unit: used to supplement missing data of numerical data in multi-source traffic data;

[0013] Anomaly detection unit: After the missing data of the road section number partition is supplemented by the missing data processing unit, a parallel clustering algorithm based on DBSCAN is used to perform distributed clustering on the data in the road section number partition on multiple nodes of the Hadoop cluster;

[0014] By calculating the local reachable density of each data point, generating the cluster center of the spatiotemporal dimension through the OPTICS algorithm, constructing the dynamic threshold range through the quartile method, and monitoring data fluctuations in real time and identifying abnormal data through the Xbar-S control chart algorithm;

[0015] Cleaning strategy decision unit: After the anomaly detection unit performs anomaly detection on the road section number partition, the data missing rate δ and the abnormal data ratio ε are calculated through statistical analysis, and the three-level cleaning mode is dynamically selected;

[0016] Preferably, in the missing data processing unit:

[0017] When data in a certain road section number partition is missing, a machine learning algorithm is used to generate predicted values for the missing data for the road section number partition to supplement the missing data.

[0018] Preferably, by calculating the local reachability density of each data point as follows:

[0019] Assume that any two data points are p and q, and calculate the spatiotemporal joint distance d between data points p and q st (p, q):

[0020] Among them, d s (p, q) is the spatial Euclidean distance between p and q, d t (p, q) is the time difference between p and q, w s d s The weight coefficient of (p, q), w t d t The weight coefficient of (p, q);

[0021] Local reachability density calculation:

[0022] For a data point p, calculate the reachable distance reach_dist of its k nearest neighbors k (p, q):

[0023] reach_dist k (p,q)=max(∈,d st (p, q));

[0024] The k nearest neighbor points are the k data points closest to the data point p;

[0025] ∈ is the neighborhood radius parameter of the DBSCAN algorithm;

[0026] Then calculate the average reachable distance avg_reach_dist(p) of the k nearest neighbors of the data point p:

[0027]

[0028] N k (p) is the set of k nearest neighbor points of data point p;

[0029] Then the local reachable density LRD(p) of data point p is:

[0030]

[0031] Preferably, w s and w t Obtained by the following method:

[0032] Calculate the Euclidean distance variance of all data points in the historical data in the spatial dimension The variance of all timestamps in the historical data

[0033]

[0034] Every fixed time window, recalculate the data in the current window and Update w s and w t The numerical value of .

[0035] Preferably, the three-level cleaning mode triggering conditions of the cleaning strategy decision unit are:

[0036] When δ<5% and ε<3%, the first-level cleaning is enabled. The first-level cleaning adopts the LSTM model based on particle swarm optimization and integrates spatiotemporal correlation to perform sequence repair.

[0037] When 5%≤δ<15% or 3%≤ε<8%, the second-level cleaning is enabled, and the second-level cleaning adopts the joint interpolation of support vector regression prediction and cubic spline interpolation;

[0038] When δ ≥ 15% or ε ≥ 8%, the third-level cleaning is enabled. The third-level cleaning directly deletes abnormal data and marks the data as missing.

[0039] Preferably, in the quality assessment module, a quality template library containing historical high-quality data features is constructed, and the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library is calculated.

[0040] If the similarity is lower than the threshold, the cleaning strategy upgrade is triggered. The cleaning strategy upgrade is:

[0041] If the current cleaning method is level one, upgrade to level two.

[0042] If the system is currently in the second level of cleaning, upgrade to the third level of cleaning.

[0043] Preferably, the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library is calculated as follows:

[0044] The cleaned data is processed through data integration, feature extraction, feature selection and data standardization to generate a data vector x, where x = (x1, x2, ..., x n ), the historical quality template generates a data vector y after data integration, feature extraction, feature selection and data standardization, y=(y1,y2,…,y n );

[0045] Calculate the spatiotemporal similarity Sim between data vector x and data vector y:

[0046]

[0047] Where n represents the number of features; i = 1, 2, ..., n; d(x, y) is the Euclidean distance between data vector x and data vector y; e is a natural constant used to construct the exponential function e d(x,y) .

[0048] Preferably, in the data collection and distributed storage module, hash partitioning is used when partitioning the multi-source traffic data according to the road section number.

[0049] The cloud computing-based traffic big data cleaning system proposed in this invention has the following beneficial technical effects:

[0050] The data missing rate and the proportion of abnormal data are calculated through statistical analysis, and the three-level cleaning mode is dynamically selected to achieve intelligent adjustment of the cleaning intensity. Through the settings in the quality assessment module, the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library is calculated, the cosine similarity and the attenuation factor based on Euclidean distance are combined, and the direction and distance information of the data are comprehensively considered. The similarity between the cleaned data and the historical quality templates is evaluated from multiple dimensions, avoiding the one-sidedness of single indicator evaluation, supporting the upgrade of the cleaning strategy, forming a closed-loop feedback for data quality improvement, and reducing the risk of misjudgment of the cleaning strategy decision-making unit.

[0051] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a principle block diagram of the system of the present invention. DETAILED DESCRIPTION

[0053] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar symbols throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention, and are not to be construed as limiting the present invention.

[0054] like Figure 1 The traffic big data cleaning system shown in the figure based on cloud computing includes:

[0055] Data acquisition and distributed storage module: used to collect multi-source traffic data in real time, including road monitoring video and image data, meteorological data, and location information data, speed data, traffic flow data, and operation trajectory data of transportation vehicles;

[0056] By matching with the map database, the corresponding road section number information is added to each multi-source traffic data, and the multi-source traffic data is partitioned by road section number and stored in the distributed file system of the cloud platform;

[0057] When partitioning multi-source traffic data by road section number, hash partitioning is used to divide the multi-source traffic data into different partitions;

[0058] Multi-level dynamic cleaning module: used to clean numerical data in multi-source traffic data;

[0059] Multi-stage dynamic cleaning module: including:

[0060] Missing data processing unit: used to supplement missing data of numerical data in multi-source traffic data;

[0061] Furthermore, in the missing data processing unit:

[0062] When data in a certain road section number partition is missing, a machine learning algorithm is used to generate the predicted value of the missing data for the road section number partition to supplement the missing data;

[0063] Anomaly detection unit: After supplementing the missing data of the road section number partition through the missing data processing unit, a parallel clustering algorithm based on DBSCAN is used to perform distributed clustering of the road section data on multiple nodes of the Hadoop cluster;

[0064] Based on the calculation of the local reachable density of each data point, the OPTICS algorithm is used to generate the cluster center of the spatiotemporal dimension. The dynamic threshold range is constructed using the quartile method. The Xbar-S control chart algorithm is used to monitor data fluctuations in real time and identify abnormal data.

[0065] DBSCAN is a density-based spatial clustering algorithm. The DBSCAN-based parallel clustering algorithm uses parallel computing technology to optimize the calculation process of the DBSCAN algorithm to improve the algorithm's operating efficiency and ability to process large-scale data.

[0066] The OPTICS algorithm is a density-based spatial clustering algorithm, which is a density clustering algorithm used for cluster analysis;

[0067] A Hadoop cluster is a distributed system consisting of multiple nodes, mainly used to store and process massive amounts of data;

[0068] The Xbar-S control chart algorithm is a commonly used tool in statistical process control to monitor the stability of production processes or data and promptly detect abnormal fluctuations in data;

[0069] Furthermore, the local reachability density of each data point is calculated as follows:

[0070] Assume that any two data points are p and q, and calculate the spatiotemporal joint distance d between data points p and q st (p, q):

[0071] Among them, d s (p, q) is the spatial Euclidean distance between p and q, d t (p, q) is the time difference between p and q, w s d s The weight coefficient of (p, q), w t d t The weight coefficient of (p, q);

[0072] Local reachability density calculation:

[0073] For a data point p, calculate the reachable distance reach_dist of its k nearest neighbors k (p, q):

[0074] reach_dist k (p,q)=max(∈,d st (p, q));

[0075] The k nearest neighbor points are the k data points closest to the data point p;

[0076] ∈ is the neighborhood radius parameter of the DBSCAN algorithm;

[0077] Then calculate the average reachable distance avg_reach_dist(p) of the k nearest neighbors of the data point p:

[0078]

[0079] N k (p) is the set of k nearest neighbor points of data point p;

[0080] Then the local reachable density LRD(p) of data point p is:

[0081]

[0082] Furthermore, w s and w t Obtained by the following method:

[0083] Calculate the Euclidean distance variance of all data points in the historical data in the spatial dimension The variance of all timestamps in the historical data

[0084]

[0085] w s and w t The proportions are inversely proportional, with the purpose of:

[0086] If the data distribution of a certain dimension has a large discrete variance, it means that the contribution of this dimension to the distance calculation may be overestimated, so its weight is reduced; for dimensions with a small variance, the data distribution is concentrated, so the weight is increased to enhance its ability to distinguish in clustering;

[0087] Every fixed time window, recalculate the data in the current window and Update w s and w t The value of

[0088] Through this dynamic calibration, the system can adapt to the spatiotemporal distribution characteristics of traffic data and ensure the accuracy of cluster center generation;

[0089] Cleaning strategy decision unit: After the anomaly detection unit performs anomaly detection on the road section number partition, the data missing rate δ and the abnormal data ratio ε are calculated through statistical analysis, and the three-level cleaning mode is dynamically selected;

[0090] Furthermore, the three-level cleaning mode triggering conditions of the cleaning strategy decision unit are:

[0091] When δ<5% and ε<3%, first-level cleaning is enabled. First-level cleaning uses an LSTM model based on particle swarm optimization and integrates spatiotemporal correlation for sequence repair. It works best when the data quality is relatively good.

[0092] LSTM stands for Long Short-Term Memory (LSTM). The LSTM model based on particle swarm optimization (PSO) combines the advantages of both algorithms and LSTM. It aims to improve the performance of LSTM models when processing complex time series data and is an existing technology.

[0093] When 5%≤δ<15% or 3%≤ε<8%, the second-level cleaning is enabled. The second-level cleaning uses SVR regression prediction and cubic spline interpolation to interpolate, and its processing capability is stronger than that of the first-level cleaning.

[0094] SVR is a support vector regression algorithm, and cubic spline interpolation is an interpolation method. The existing technology is to use SVR regression prediction and cubic spline interpolation to interpolate.

[0095] When δ ≥ 15% or ε ≥ 8%, the third-level cleaning is enabled. The third-level cleaning directly deletes abnormal data and marks the position of missing data. By deleting abnormal data, it can avoid interference with subsequent analysis;

[0096] Quality assessment module: Build a quality template library containing historical high-quality data features, calculate the spatiotemporal similarity between cleaned data and historical quality templates in the quality template library,

[0097] If the similarity is lower than the threshold, the cleaning strategy upgrade is triggered. If the current cleaning strategy is level one, it is upgraded to level two.

[0098] If the current cleaning is at level 2, upgrade to level 3;

[0099] Furthermore, the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library is calculated as follows:

[0100] The cleaned data is processed through data integration, feature extraction, feature selection and data standardization to generate a data vector x, where x = (x1, x2, ..., x n ), the historical quality template generates a data vector y after data integration, feature extraction, feature selection and data standardization, y=(y1,y2,…,y n );

[0101] Calculate the spatiotemporal similarity Sim between data vector x and data vector y:

[0102]

[0103] n represents the number of features;

[0104] i=1, 2, ..., n;

[0105] d(x, y) is the Euclidean distance between data vector x and data vector y;

[0106] e is a natural constant used to construct the exponential function e d(x,y) ;

[0107] By combining cosine similarity and the attenuation factor based on Euclidean distance, the direction and distance information of the data are comprehensively considered, and the similarity between the cleaned data and the historical quality template is evaluated from multiple dimensions, avoiding the one-sidedness of single indicator evaluation.

[0108] In the context of traffic big data, data has spatiotemporal characteristics. Cosine similarity can capture the changing trends of data in time series and the similarity of distribution patterns in spatial characteristics, while Euclidean distance can reflect the absolute differences in data values at different time points or spatial locations. Therefore, the composite evaluation function can better adapt to the spatiotemporal characteristics of traffic big data, more accurately evaluate data quality, and improve the accuracy of similarity measurement;

[0109] Euclidean distance-based decay factor The nonlinear property of the composite evaluation function can penalize data points with large distances, reducing their impact on the overall similarity assessment. This makes the composite evaluation function more robust when processing data containing outliers, reducing the interference of outliers on the evaluation results and improving the accuracy of the evaluation.

[0110] The data missing rate and the proportion of abnormal data are calculated through statistical analysis, and the three-level cleaning mode is dynamically selected to achieve intelligent adjustment of the cleaning intensity. Through the settings in the quality assessment module, the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library is calculated, the cosine similarity and the attenuation factor based on Euclidean distance are combined, and the direction and distance information of the data are comprehensively considered. The similarity between the cleaned data and the historical quality templates is evaluated from multiple dimensions, avoiding the one-sidedness of single indicator evaluation, supporting the upgrade of the cleaning strategy, forming a closed-loop feedback for data quality improvement, and reducing the risk of misjudgment of the cleaning strategy decision-making unit.

[0111] A traffic big data cleaning method based on cloud computing includes the following steps:

[0112] S1. Real-time collection of multi-source traffic data, including road monitoring video and image data, meteorological data, and location information data, speed data, traffic flow data, and operation trajectory data of transportation vehicles;

[0113] S2. By matching with the map database, corresponding road section number information is added to each piece of multi-source traffic data, and the multi-source traffic data is partitioned by road section number and stored in the distributed file system of the cloud platform;

[0114] When data in a certain road section number partition is missing, a machine learning algorithm is used to generate the predicted value of the missing data for the road section number partition to supplement the missing data;

[0115] S3, after supplementing the missing data of the road section number partition through S2, a parallel clustering algorithm based on DBSCAN is used to perform distributed clustering on the data in the road section number partition on multiple nodes of the Hadoop cluster;

[0116] By calculating the local reachable density of each data point, generating the cluster center of the spatiotemporal dimension through the OPTICS algorithm, constructing the dynamic threshold range through the quartile method, and monitoring data fluctuations in real time and identifying abnormal data through the Xbar-S control chart algorithm;

[0117] S4: After performing anomaly detection on the road section number partitions through S3, the data missing rate and abnormal data ratio are calculated through statistical analysis, and the three-level cleaning mode is dynamically selected;

[0118] Build a quality template library containing historical high-quality data features, calculate the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library,

[0119] S5. If the similarity is lower than the threshold, the cleaning strategy upgrade is triggered. The cleaning strategy upgrade is:

[0120] If the current cleaning method is level one, upgrade to level two.

[0121] If the system is currently in the second level of cleaning, upgrade to the third level of cleaning.

[0122] Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0123] In the embodiments provided by the present invention, it should be understood that the disclosed systems or methods can be implemented in other ways. For example, the embodiments of the invention described above are merely illustrative. For example, the division of modules is only a logical function division, and other division methods may be used in actual implementation.

[0124] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected to achieve the objectives of this embodiment based on actual needs.

[0125] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing module, each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or hardware plus software functional modules.

[0126] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the basic characteristics of the present invention.

[0127] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A traffic big data cleaning system based on cloud computing, characterized by: include: Data collection and distributed storage module: used to collect numerical data from multi-source traffic data in real time. The numerical data in multi-source traffic data includes location information data, speed data, traffic flow data and operation trajectory data of transportation vehicles; By matching with the map database, the corresponding road section number information is added to each multi-source traffic data, and the multi-source traffic data is partitioned by road section number and stored in the distributed file system of the cloud platform; Multi-level dynamic cleaning module: used to clean numerical data in multi-source traffic data; Quality assessment module: used to perform quality assessment on data cleaned by the multi-level dynamic cleaning module.

2. The cloud computing-based traffic big data cleaning system according to claim 1 is characterized in that: The multi-stage dynamic cleaning module includes: Missing data processing unit: used to supplement missing data of numerical data in multi-source traffic data; Anomaly detection unit: After the missing data of the road section number partition is supplemented by the missing data processing unit, a parallel clustering algorithm based on DBSCAN is used to perform distributed clustering on the data in the road section number partition on multiple nodes of the Hadoop cluster; By calculating the local reachable density of each data point, generating the cluster center of the spatiotemporal dimension through the OPTICS algorithm, constructing the dynamic threshold range through the quartile method, and monitoring data fluctuations in real time and identifying abnormal data through the Xbar-S control chart algorithm; Cleaning strategy decision unit: After the road section number partition is detected by the anomaly detection unit, the data missing rate δ and the abnormal data proportion ε are calculated through statistical analysis, and the three-level cleaning mode is dynamically selected.

3. The cloud computing-based traffic big data cleaning system according to claim 2 is characterized in that: In the missing data processing unit: When data in a certain road section number partition is missing, a machine learning algorithm is used to generate predicted values for the missing data for the road section number partition to supplement the missing data.

4. The cloud computing-based traffic big data cleaning system according to claim 2 is characterized in that: By calculating the local reachability density of each data point, as follows: Assume that any two data points are p and q, and calculate the spatiotemporal joint distance d between data points p and q st (p, q): Among them, d s (p, q) is the spatial Euclidean distance between p and q, d t (p, q) is the time difference between p and q, w s d s The weight coefficient of (p, q), w t d t The weight coefficient of (p, q); Local reachability density calculation: For a data point p, calculate the reachable distance reach_dist of its k nearest neighbors k (p,q): reach_dist k (p,q)=max(∈,d st (p,q); The k nearest neighbor points are the k data points closest to the data point p; ∈ is the neighborhood radius parameter of the DBSCAN algorithm; Then calculate the average reachable distance avg_reach_dist(p) of the k nearest neighbors of the data point p: N k (p) is the set of k nearest neighbor points of data point p; Then the local reachable density LRD(p) of data point p is:

5. The cloud computing-based traffic big data cleaning system according to claim 4 is characterized in that: w s and w t Obtained by the following method: Calculate the Euclidean distance variance of all data points in the historical data in the spatial dimension The variance of all timestamps in the historical data Every fixed time window, recalculate the data in the current window and Update w s and w t The numerical value of .

6. The cloud computing-based traffic big data cleaning system according to claim 2 is characterized in that: The three-level cleaning mode triggering conditions of the cleaning strategy decision unit are: When δ<5% and ε<3%, the first-level cleaning is enabled. The first-level cleaning adopts the LSTM model based on particle swarm optimization and integrates spatiotemporal correlation to perform sequence repair. When 5%≤δ<15% or 3%≤ε<8%, secondary cleaning is enabled, and secondary cleaning uses support vector regression prediction and cubic spline interpolation; When δ ≥ 15% or ε ≥ 8%, the third-level cleaning is enabled. The third-level cleaning directly deletes abnormal data and marks the data as missing.

7. The cloud computing-based traffic big data cleaning system according to claim 6 is characterized in that: In the quality assessment module, a quality template library containing historical high-quality data features is constructed, and the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library is calculated. If the similarity is lower than the threshold, the cleaning strategy upgrade is triggered. The cleaning strategy upgrade is: If the current cleaning method is level one, upgrade to level two. If the system is currently in the second level of cleaning, upgrade to the third level of cleaning.

8. The cloud computing-based traffic big data cleaning system according to claim 7 is characterized in that: Calculate the spatiotemporal similarity between the cleaned data and the historical quality templates in the quality template library as follows: The cleaned data is processed through data integration, feature extraction, feature selection and data standardization to generate a data vector x, where x = (x1, x2, ..., x n ), the historical quality template generates a data vector y after data integration, feature extraction, feature selection and data standardization, y=(y1,y2,…,y n ); Calculate the spatiotemporal similarity Sim between data vector x and data vector y; Where n represents the number of features; i = 1, 2, ..., n; d(x, y) is the Euclidean distance between data vector x and data vector y; e is a natural constant used to construct the exponential function e d(x,y) .

Citation Information

Patent Citations

  • A method for cleaning traffic big data based on a cloud computing framework

    CN106202335B

  • Traffic big data cleaning method based on cloud computing framework

    CN106202335A

  • Traffic big data cleaning method, device and equipment and readable storage medium

    CN114281808A

  • State monitoring system and method for road and bridge anti-collision guardrail

    CN118071114A

  • Medical examination result automatic interpretation and diagnosis support system

    CN118507033A

Cited By

  • Large-scale trajectory data cleaning and traffic situation intelligent evaluation system

    CN121256621A

  • Large-scale trajectory data cleaning and intelligent traffic situation assessment system

    CN121256621B