A Quality Assessment Method for Spatiotemporal Trajectory Big Data
By clustering and redistributing trajectory data, combined with multi-dimensional constraint inspection, the efficiency and accuracy of trajectory data quality evaluation in the prior art are solved, and efficient quality evaluation of batch and streaming trajectory data is achieved.
Patent Information
- Application Number
- CN202510347554.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The prior art is difficult to effectively capture spatiotemporal patterns, group behaviors and geographical environment constraints in trajectory data, cannot process streaming data, and is inefficient, which cannot meet the needs of trajectory spatiotemporal data quality assessment.
The clustering and redistribution methods are used to convert the trajectory data set into a grid-based representation form, and the trajectory subset is obtained through sampling, combining timestamps, ranges, smoothness, road network constraints, etc. to check the trajectory quality, and support offline and online evaluation of batch and streaming data.
It realizes constrained inspection of the validity, completeness, consistency and fairness of trajectory data, improves the efficiency and accuracy of trajectory data quality evaluation, and is suitable for batch and streaming data.
Smart Images

Figure CN119862401B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data management and quality assessment, and particularly relates to a quality assessment method for spatio-temporal trajectory big data. Background Art
[0002] With the popularization of GPS (Global Positioning System) and the wide use of related services, trajectory data plays an important role in multiple fields such as transportation and smart cities. However, due to problems such as inaccurate GPS measurements, low sampling rates, and transmission interruptions, trajectory data often contains errors, resulting in a decline in the quality of trajectory data, which in turn has a negative impact on downstream services. Therefore, evaluating the quality of trajectory data is a crucial but cumbersome task, which provides guidance for subsequent data cleaning and analysis.
[0003] To solve related problems, some existing studies involve general data quality assessment, which can be mainly divided into two categories: task-independent and task-related. Among them, task-independent data quality assessment, such as the literature [Yunxiang Su, Yikun Gong, Shaoxu Song. 2023. Validity of Time Series Data. ACM Transactions on Database Systems. 2023.1.1, Volume 85:1-85:26], uses data characteristics to quantify various quality dimensions; task-related data quality assessment, such as the literature [Cedric Lengyel, Luka Rimanić, Luka Kolar, Wentao Wu, Ce Zhang. 2023. Automatic Feasibility Study of Machine Learning Based on Data Quality Analysis: A Case Study of Label Noise. Proceedings of the International Conference on Data Engineering. pp. 218-231], considers data quality in the context of specific tasks, especially in the field of machine learning.
[0004] However, existing technologies mainly focus on general data quality assessment problems and are not good at handling data quality assessment tasks specific to the trajectory level. First, the trajectory data contains spatio-temporal dependence characteristics. Existing technologies only focus on patterns in general data and lack the ability to effectively capture violations of spatio-temporal patterns in trajectory data. Second, trajectory data usually exhibits specific group behaviors, such as morning and evening rush hours. Existing technologies only focus on detecting errors in a single univariate time series and lack the ability to consider interdependencies. In addition, trajectory data is restricted by specific geographical environments. For example, capturing the trajectory of a car's movement in an urban area must be consistent with the urban road network. Existing technologies ignore auxiliary information such as road networks and cannot capture the potential relationship between the terrain context and the trajectory. Moreover, trajectory data is often generated in the form of data streams and used in real time. Existing technologies only focus on static data and cannot be effectively extended to handle streaming data. Although some technologies support incremental quality assessment, they still face efficiency challenges when dealing with complex constraint checks. The above problems make it difficult for existing data quality assessment technologies to be applied to real trajectory spatio-temporal data quality assessment scenarios. Summary of the Invention
[0005] In view of the above, the present invention provides a quality assessment method for spatio-temporal trajectory big data, which supports constraint checks covering four dimensions (validity, integrity, consistency, and fairness), supports offline and online evaluation of batch historical trajectory data and streaming real-time trajectory data, and adopts an evaluation optimization strategy to improve evaluation efficiency.
[0006] A quality assessment method for spatio-temporal trajectory big data includes the following steps:
[0007] (1) Obtain a trajectory dataset, perform clustering and reallocation on the trajectories in the trajectory dataset to obtain multiple clusters with a uniform number;
[0008] (2) Selectively sample a trajectory subset from all clusters;
[0009] (3) Perform quality assessment on the trajectory subset from the dimension of data validity, including checking the correctness of trajectory point timestamp constraints and range constraints;
[0010] (4) Perform quality assessment on the trajectory subset from the dimension of data integrity, including checking for missing trajectory points and missing values of trajectory points;
[0011] (5) Perform quality assessment on the trajectory subset from the dimension of data consistency, including checking for trajectory smoothness constraints, trajectory length constraints, abnormal trajectories, abnormal trajectory stop points, and trajectory road network constraints;
[0012] (6) Conduct quality assessment on the trajectory subset from the dimension of data fairness, including calculating the spatial density of trajectory points within a specific area and the temporal density of trajectory points within a specific time period;
[0013] (7) According to the statistics in steps (3) to (6) of the number and proportion of detected trajectories and trajectory points, together with the spatial density and temporal density calculated in step (7), output them as the final quality assessment results.
[0014] Further, the trajectory dataset is collected through GPS devices to obtain a large number of trajectories. Each trajectory is represented by a number of trajectory points, and each trajectory point contains information such as index, timestamp, longitude, and latitude; the trajectory dataset is a batch trajectory dataset or a streaming trajectory dataset. The batch trajectory dataset stores historical trajectory data in an offline storage manner, and the streaming trajectory data stores real-time trajectory data in an online storage manner.
[0015] Further, the specific process of clustering and reassigning the trajectories in the trajectory dataset in step (1) is as follows:
[0016] 1.1 First, divide the entire two-dimensional geographical space into several grid cells of equal size, and then replace each trajectory point in the trajectory with the grid cell it belongs to, thereby converting the trajectories in the trajectory dataset into a grid-based representation form;
[0017] 1.2 Sort the trajectories in the trajectory dataset in descending order of length, take the first trajectory, that is, the trajectory with the largest length, and initialize a cluster with this trajectory as the cluster center;
[0018] 1.3 Take the next trajectory in order T G , compare it with the cluster centers of all clusters. If there is a certain cluster whose cluster center coincides with T G by k grid cells, then add T G to this cluster. k is a set threshold. At the same time, if the length of T G is greater than the length of the cluster center of this cluster, then use T G as the new cluster center of this cluster; if there is no such cluster, then create a new cluster with T G as the cluster center;
[0019] 1.4 Cluster the trajectories in the trajectory dataset into multiple clusters in the same way as in step 1.3, and randomly select some trajectories from the large-sized clusters with more than the average number of trajectories and assign them to the small-sized clusters with less than the average number of trajectories to ensure an even number of trajectories in each cluster.
[0020] Further, the specific implementation of step (2) is as follows:
[0021] 2.1 Randomly select a cluster, divide the trajectories in this cluster into original trajectories and assigned trajectories according to their sources, and determine the proportion of these two types of trajectories in the cluster. The original trajectories are the trajectories obtained through clustering, and the assigned trajectories are the trajectories obtained through reallocation from other clusters;
[0022] 2.2 Sample a random number from a uniform distribution between 0 and 1. If this random number is greater than a , then randomly select one trajectory from the original trajectories in this cluster and include it in the trajectory subset. Otherwise, randomly select one trajectory from the assigned trajectories in this cluster and include it in the trajectory subset. a is the proportion of the original trajectories in the cluster;
[0023] 2.3 Repeatedly execute steps 2.1 to 2.2 until the number of trajectories in the trajectory subset reaches the set quantity requirement. This sampling method can ensure that the quality evaluation result on the trajectory subset is approximately the same as that on the complete trajectory dataset, reducing the computational cost.
[0024] Further, the timestamp constraint in step (3) requires that the timestamps of the trajectory points on the same trajectory increase monotonically with time, and the range constraint requires that the longitude and latitude of each trajectory point be within the range of the longitude and latitude limits of the specified area.
[0025] Further, for any two consecutive trajectory points on the trajectory in step (4), if the time interval t between these two points is greater than 2Δ t , then it is determined that there are missing points on this trajectory, and the number of missing points between these two trajectory points is , and Δ t is the sampling period of this trajectory, that is, the minimum value of the time intervals between all consecutive two points on this trajectory; if any trajectory point has no longitude value, latitude value, or timestamp, then it is determined that there is a missing value for this trajectory point.
[0026] Further, the trajectory smoothness constraint in step (5) detects the position deviation through a Kalman filter, and requires that there be no sudden position movement of the trajectory points in the trajectory; the trajectory length constraint requires that the cumulative value of the Haversine distances between all consecutive two points in the trajectory ≥ L th , L this the trajectory length threshold and is set at three times the standard deviation of the mean trajectory length;
[0027] The judgment criteria for abnormal trajectories are: for any trajectory in the trajectory subset T , if the trajectory T The number of outliers in is greater than a given threshold ρ , then determine T is an abnormal trajectory; for the trajectory T Any track point p ,like p The number of neighbor points of is less than the given threshold η , then it is determined p is an outlier point; for any trajectory point on other trajectories in the trajectory subset q , if dist( p , q )≤ σ , then it is determined q for p Neighboring points, dist( p , q )for p and q Haversine distance, σ is a given distance threshold;
[0028] The judgment criteria of abnormal trajectory stop points are as follows: first, all possible stop points are identified by using the TrajDBSCAN (Trajectory Density Clustering Spatial Noise Application) algorithm of a small neighborhood for all trajectories in the trajectory subset, and then cluster analysis is performed on the identified stop points, and the stop points that do not belong to any cluster are determined as abnormal trajectory stop points;
[0029] The trajectory network constraint requires that the trajectory must match the path in the road network. The OHMM (based on online hidden Markov model) network matching algorithm is used to find the correspondence between each trajectory point in the trajectory and the road section in the road network. If the corresponding relationship does not exist, the corresponding trajectory point and its trajectory are judged to violate the trajectory network constraint.
[0030] Furthermore, in step (6), for a specific area Q S , the spatial density of trajectory points in this area ,in N is the total number of trajectory points in the trajectory subset, N ( Q S ) is the trajectory subset located in a specific area Q S The number of trajectory points in Area is the spatial distribution area of the trajectory subset, Area (Q S ) is the area of a specific region Q S ; for a specific time period Q t , the time density of the trajectory points within this time period , N ( Q t ) is the number of trajectory points in the trajectory subset within a specific time period Q t , Time is the overall time span of the trajectory subset, Time ( Q t ) is the time span of a specific time period Q t .
[0031] A computer device includes a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the computer program to implement the above-mentioned quality assessment method for spatio-temporal trajectory big data.
[0032] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned quality assessment method for spatio-temporal trajectory big data.
[0033] The present invention designs a quality assessment method for spatio-temporal trajectory big data, which supports constraint checking covering four dimensions (validity, integrity, consistency, and fairness), supports offline and online evaluation of batch historical trajectory data and streaming real-time trajectory data, and adopts an evaluation optimization strategy to improve evaluation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a schematic flowchart of the quality assessment method for spatio-temporal trajectory big data according to the present invention.
[0035] Figure 2 is a schematic flowchart of the operation of clustering and reassigning trajectories according to the present invention.
[0036] Figure 3 is a schematic diagram of trajectory data reallocation and representative trajectory selection in an embodiment of the present invention.
[0037] Figure 4 is a schematic flowchart of the operation of sampling trajectories in a cluster to obtain a trajectory subset according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] To describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0039] like Figure 1 As shown, the quality assessment method for spatiotemporal trajectory big data of the present invention is applied in a terminal and includes the following steps:
[0040] Step S11: Obtain a trajectory data set, which consists of a number of trajectories, each of which is represented by a number of trajectory points, wherein each trajectory point contains a trajectory index, a timestamp, a longitude, and a latitude; for the trajectory data stored in batches, an offline storage method is used for storage, and for the trajectory data flowing in in real time, an online storage method is used for storage.
[0041] Specifically, the trajectory dataset is mainly obtained through GPS and consists of several trajectories. Each trajectory consists of several GPS sampling points, which become a data record, containing information such as the longitude, latitude and timestamp of the sampling point, that is, . Historical batch trajectory data refers to trajectory data that has been collected and stored. This data can be processed periodically or once. For this type of data, an offline storage method can be used, that is, the data is stored in a database or data warehouse, and then analyzed and processed through batch jobs. The offline storage method allows complex queries and analysis of large amounts of data, but is not suitable for scenarios that require real-time response. Real-time streaming trajectory data refers to data collected continuously from mobile objects. This data needs to be processed immediately to support real-time applications. For streaming data, an online storage method is used, that is, the data is stored in memory or a high-speed storage system and can be processed and analyzed in real time. The online storage method supports fast data access and processing and is suitable for applications that require instant feedback. For offline storage, traditional relational database management systems (RDBMS) or object-oriented databases such as PostgreSQL (Postgresql) and MySQL (My Structured Query Language) can be used. For online storage, you might want to use an in-memory database like Redis (a remote dictionary service), or a stream processing framework like Apache Kafka, Apache Flink, or Spark Streaming.
[0042] Step S12: Cluster and redistribute the trajectories to obtain a uniform number of trajectory clusters. The specific operation process is as follows: Figure 2 As shown:
[0043] S21: Convert each original trajectory data into a grid representation form. This process involves dividing the two-dimensional geographic space into a number of grid cells of equal size, and representing each trajectory point with the grid cell identifier to which the trajectory point belongs.
[0044] Specifically, the trajectory represented in the free geospatial space is transformed into the grid space for representation. Given a two-dimensional free space, we construct a grid index by dividing the space into 2 θ × 2 θ equal-sized grid cells G where 2 θ is the grid resolution. By replacing each point in the trajectory with the ID of the grid cell it belongs to, the free geospatial trajectory is transformed into a grid space trajectory. After applying the above method, the trajectory is simplified to a discrete sequence of grid cells.
[0045] S22: Sort the grid-based trajectories in descending order of trajectory length and perform the clustering process: Traverse each trajectory to check if the overlapping grid cells of the trajectory with the centroid trajectory of a certain trajectory cluster exceed a certain threshold. If so, add the trajectory to the corresponding cluster; otherwise, form a new cluster with the trajectory as a centroid.
[0046] Specifically, given a grid space trajectory dataset, first sort the trajectories in descending order of length, and then initialize an empty list to store the centroids of each cluster. In the actual clustering process, the longest trajectory in the cluster is regarded as the centroid. Traverse the sorted grid space trajectories. For the currently visited trajectory T G , check if it has at least k grid cells intersecting with any current centroid. If T G it has at least k grid cells overlapping with the centroid c , add it to the corresponding cluster with c as the centroid and determine if the centroid of the cluster needs to be modified; otherwise, initialize a new cluster with C as the centroid T G . This process continues until all trajectories have been visited. C new
[0047] S23: Randomly select some trajectories from the large-sized clusters with more than the average number of trajectories and assign them to the small-sized clusters with fewer than the average number of trajectories to ensure that the number of trajectories in each cluster is the same.
[0048] Specifically, the size of each cluster data volume is normalized by where | C i | represents the number of trajectories in cluster i , K represents the number of clusters, and the denominator represents the total number of trajectories. After normalization, the average number of trajectories in all clusters is 1;Figure 3 In a specific embodiment shown, the normalized sizes of each cluster are 0.4, 1.2, 0.8, and 1.6 respectively. The reallocation method first creates two queues Q >1 and Q <1 , Q >1 stores the cluster IDs with a data volume greater than 1, while Q <1 stores the cluster IDs with a data volume less than 1, and initializes an empty dictionary A to record the allocation relationship between the data distributor and the data receiver. When both queues are not empty, the algorithm takes out a cluster Q >1 from l , and at the same time takes out a cluster Q <1 from s , then records in A that s receives the trajectory data from l to make its own data volume satisfy 1, and at the same time updates the size of cluster l . According to the updated size, it is judged whether the cluster l needs to be added to Q >1 again. The process continues until the sizes of all clusters are balanced to 1 or the queue becomes empty, and finally clusters with uniform sizes are obtained. Figure 3 In
[0049] Step S13: Selectively sample the trajectory data in each uniform trajectory cluster to obtain a trajectory subset, and this sampling ensures that the quality evaluation result of the complete trajectory data set is approximately the same as the quality evaluation result on this trajectory subset. The specific operation process is as Figure 4 shown:
[0050] S31: Randomly select a specific cluster, and sample a trajectory from this cluster and include it in the trajectory subset.
[0051] Specifically, after the clustering process, n clusters are obtained. Use the random number generation algorithm to generate a random integer i ∈[1, n , and subsequently sample the cluster with ID i .
[0052] S32: The trajectories in the cluster can be divided into an original part (from the current cluster) and an allocated part (from other clusters) according to the source. The proportion of the number of trajectories in the original part to the total number of trajectories in the cluster isp (The ratio of the allocated part is 1 - p ), sample a random number from a uniform distribution from 0 to 1. If the random number is greater than p then sample a trajectory from the original part, otherwise sample a trajectory from the allocated part.
[0053] Specifically, the Alias sampling process first generates a random number uniformly distributed between 0 and 1 r , if r is less than the normalized size of the cluster with the selected ID i which is d i , then sample a trajectory from the cluster C i , otherwise sample a trajectory from the corresponding cluster A in the data distributor - receiver relationship C Ai . This process is repeated until representative trajectories meeting the quantity requirements are sampled; this method ensures that the sampled representative trajectories are consistent with the overall trajectory data distribution. In a specific embodiment as shown in Figure 3 , in the first case, cluster 1 is selected because the random number r = 0.3 is less than the normalized size 0.4 of cluster 1, so a trajectory is sampled from the yellow part (original data) in cluster 1; in the second case, cluster 3 is selected because the random number r = 0.9 is greater than the normalized size 0.8 of cluster 3, so a trajectory is sampled from the green part (allocated data) in cluster 3. This redirection mechanism ensures that even after trajectory reallocation, both the original and allocated parts of the data in the cluster can be appropriately sampled.
[0054] S33: Repeat the above sampling process until the number of trajectories in the trajectory subset reaches a preset number.
[0055] Step S14: Evaluate the quality of the trajectory subset from the perspective of data validity, including checks on the correctness of individual trajectory points such as timestamp constraints and range constraints.
[0056] Specifically, validity refers to the degree to which trajectory data conforms to the basic spatio - temporal pattern, and it imposes timestamp and range constraints on individual points in the trajectory. The timestamp constraint ensures that the timestamps of trajectory points increase monotonically over time, defined as , where p represents a trajectory point, represented by the triple ( x , y , t ). The range constraint limits the position range of GPS points, defined as: , where xmin and x max represent the minimum and maximum longitudes of the considered area respectively, y min and y max represent the minimum and maximum latitudes of the considered area respectively. If the area range is not specified, it is determined by the value range of the longitude and latitude themselves. In both the offline evaluation and online evaluation settings, the GPS points in the dataset are traversed to check the timestamp constraint and the range constraint.
[0057] Step S15: Evaluate the quality of the trajectory subset from the perspective of data integrity, including the inspection of missing data such as missing points and missing values in the trajectory.
[0058] Specifically, integrity refers to the completeness and information richness of the trajectory data, considering two cases: missing points in the trajectory and missing values of trajectory points. A missing point in the trajectory is defined as , where Δ t represents the sampling period of the trajectory. If the time interval between two consecutive points is abnormally long, there is a missing point. The number of missing points in the trajectory measures the sampling stability of the trajectory data. The more missing points there are, the more uneven the data sampling is. The sampling period of a trajectory is represented by the time interval with the highest frequency among all consecutive GPS points in the trajectory. In both the offline evaluation and online evaluation settings, the GPS points in the dataset are traversed to count the missing points in the trajectory. For any two consecutive points p i+1 and p i , the number of missing points between them is estimated as . A missing value of a trajectory point is defined as . If a GPS point has no longitude value, latitude value, or timestamp, we consider it a missing value. The missing values of trajectory points measure the missing situation of the longitude, latitude, and timestamp data of the trajectory points. In both the offline evaluation and online evaluation settings, the GPS points in the dataset are traversed to count the missing values of the trajectory points.
[0059] Step S16: Evaluate the quality of the trajectory subset from the perspective of data consistency, including the inspection of violations of the association relationships between trajectory points or between trajectories, such as trajectory smoothness constraints, trajectory length constraints, abnormal trajectories, abnormal trajectory stop points, and trajectory road network constraints.
[0060] Specifically, consistency refers to the degree to which the trajectory data conforms to a set of semantic rules, considering five cases: trajectory smoothness constraints, trajectory length constraints, abnormal trajectories, abnormal trajectory stop points, and trajectory road network constraints.
[0061] The trajectory smoothness constraint restricts the position change of trajectory points, requiring that there be no sudden position movement of trajectory points in the trajectory data; a Kalman filter is deployed in both the offline evaluation and online evaluation settings to detect position deviations.
[0062] The trajectory length constraint is defined as , where dist is the distance function that measures the Haversine distance between two consecutive trajectory points, L th is the trajectory length threshold, and based on the three-sigma rule, the trajectory length threshold L th is defined as three times the standard deviation of the average trajectory length. In the offline evaluation setting, the trajectories in the dataset are traversed to check the trajectory length constraint; in the online evaluation setting, only the GPS points in the sliding window are accessible. If there are no GPS points flowing in within k consecutive time windows, the trajectory stream is considered to have ended, and the trajectory length constraint is checked.
[0063] The definition of an abnormal trajectory is as follows: The neighbor of the trajectory point T from trajectory p i is defined as , where σ is a given distance threshold. If p i has a number of neighbors less than or equal to the given quantity threshold, i.e., , then p i is considered an outlier. If the number of outliers in a trajectory exceeds the given quantity threshold ρ , then the trajectory is an abnormal trajectory. In the offline evaluation setting, TRAOD (Trajectory Outlier Detection) based on the above idea is used as the abnormal trajectory detection method. This method first segments the trajectory and then detects abnormal sub-trajectory segments based on distance; in the online evaluation setting, the extended method PTMDS (Probabilistic Trajectory Matching and Detection System) of TRAOD on streaming trajectory data is used as the abnormal trajectory detection method.
[0064] A stop point refers to the area where a moving object temporarily stops. Stop points with specific geographical significance in different trajectories (such as service areas) can be well clustered, but abnormal stop points caused by unexpected situations (such as equipment failures) often appear randomly. An abnormal trajectory stop point is defined as an outlier data point (noise) generated after stop point clustering, and an abnormal trajectory stop point is a stop behavior caused by an abnormal situation. In the offline evaluation setting, first apply the TrajDBSCAN method with a small neighborhood to all trajectories in the trajectory dataset to identify all possible stop points, and then perform DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering on the identified stop points to discover abnormal trajectory stop points that do not belong to any cluster; in the online evaluation setting, use the dynamic DBSCAN algorithm to replace the TrajDBSCAN method in the offline evaluation setting to adapt to streaming trajectory data.
[0065] Given a road network represented in the form of a directed graph G =( V , E ), where the node v ∈ V represents an intersection or the end of a road, and the edge e ∈ E represents a section of road, and the path P is a sequence of connected road segments, that is . The trajectory road network constraint stipulates that each trajectory T must match a path G in P . The trajectory road network constraint measures the degree to which a trajectory follows the actual road network. In both the offline evaluation and online evaluation settings, the OHMM (Online Hidden Markov Model)-based road network matching algorithm is used to correspond each trajectory point to a road segment in the road network. Once it is found that there is no corresponding relationship between a certain trajectory point and a road segment, then this point and the trajectory to which it belongs violate the road network constraint, and the trajectory point and the corresponding trajectory violate the trajectory road network constraint.
[0066] Step S17: Evaluate the quality of the trajectory subset from the perspective of data fairness, including calculations on data distribution such as spatial density and temporal density.
[0067] Specifically, fairness refers to the degree of biased information in the trajectory dataset, which will affect the accuracy of machine learning results in subsequent applications. Fairness considers two aspects: spatial density and temporal density. The spatial density is defined as , where Q S is a rectangular area. is the set of trajectory points within in the dataset Q S . is the total number of trajectory points in the dataset, Area is the spatial distribution area of the dataset, Area ( Q S ) is Q S the spatial distribution area. The spatial density compares the density of GPS points within a specific area with the overall density, measuring the data deviation in the spatial dimension. In both the offline evaluation and online evaluation settings, a quadtree is constructed for all GPS points in the dataset to accelerate range queries through breadth-first search, so as to find . The time density is defined as , where Q t = t min , t max is the time span specified by the user, is the set of trajectory points within Q t the time span, is the overall time span of the dataset. The time density compares the density of GPS points within a specific time period with the overall density, measuring the data deviation in the time dimension. In both the offline evaluation and online evaluation settings, we construct a segment tree for each leaf node of the quadtree to efficiently perform time period queries; to find GPS points within a given time period, we traverse and prune the segment tree through breadth-first search.
[0068] Step S18: Output the count of violated metrics as the evaluation result for the given trajectory dataset.
[0069] Specifically, according to steps S14, S15, S16, and S17, the quality of a representative data subset of the trajectory dataset is evaluated from four data dimensions: data validity, integrity, consistency, and fairness. Finally, the number and proportion of trajectories and trajectory points that violate the constraints in each data dimension are output respectively. The evaluation result of the representative data subset can effectively reflect the quality of the original trajectory dataset.
[0070] The above description of the embodiments is for the convenience of those of ordinary skill in the art to understand and apply the present invention. Those who are familiar with the technology in the art can obviously make various modifications to the above embodiments easily and apply the general principles described herein to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and the improvements and modifications made by those skilled in the art according to the disclosure of the present invention should be within the protection scope of the present invention.
Claims
1. A quality assessment method for spatio-temporal trajectory big data, comprising the following steps: (1) Obtain a trajectory dataset, cluster and reassign the trajectories in the trajectory dataset to obtain multiple clusters with a uniform number; the trajectory dataset collects a large number of trajectories through a GPS device, each trajectory is represented by a number of trajectory points, and each trajectory point contains information such as an index, a timestamp, longitude, and latitude; The trajectory dataset is a batch trajectory dataset or a streaming trajectory dataset. The batch trajectory dataset stores historical trajectory data in an offline storage manner, and the streaming trajectory dataset stores real-time trajectory data in an online storage manner; The specific process of clustering and reassigning the trajectories in the trajectory dataset is as follows: 1.1 First, divide the entire two-dimensional geographical space into several grid cells of equal size, and then replace each trajectory point in the trajectory with the grid cell it belongs to, so as to convert the trajectories in the trajectory dataset into a grid-based representation form; 1.2 Sort the trajectories in the trajectory dataset in descending order of length, take the first trajectory, that is, the trajectory with the largest length, and initialize a cluster with this trajectory as the cluster center; 1.3 Take the next trajectory T in sequence G , compare it with the cluster centers of all clusters. If there exists a certain cluster whose cluster center coincides with T G in k grid cells, then add T G to this cluster. k is a set threshold. At the same time, if the length of T G is greater than the length of the cluster center of this cluster, then use T G as the new cluster center of this cluster; if there is no such cluster, then create a new cluster with T G as the cluster center; 1.4 Cluster the trajectories in the trajectory dataset into multiple clusters in the manner of step 1.3, and randomly select some trajectories from the large-sized clusters with more than the average number of trajectories and assign them to the small-sized clusters with less than the average number of trajectories to ensure that the number of trajectories in each cluster is uniform; (2) Selectively obtain a trajectory subset by sampling from all clusters, and the specific implementation method is as follows: 2.1 Randomly select a cluster, divide the trajectories in the cluster into original trajectories and assigned trajectories according to the source, and determine the proportion of these two types of trajectories in the cluster. The original trajectories are the trajectories obtained by clustering, and the assigned trajectories are the trajectories obtained by reassigning from other clusters; 2.2 Sample a random number from a uniform distribution from 0 to 1. If the random number is greater than a, randomly select a trajectory from the original trajectories of the cluster and include it in the trajectory subset. Otherwise, randomly select a trajectory from the assigned trajectories of the cluster and include it in the trajectory subset, where a is the proportion of the original trajectories in the cluster; 2.3 Repeatedly execute steps 2.1 to 2.2 until the number of trajectories in the trajectory subset reaches the set quantity requirement; (3) Conduct a quality assessment of the trajectory subset from the dimension of data validity, including checking the correctness of the timestamp constraint and range constraint of the trajectory points. The timestamp constraint requires that the timestamps of the trajectory points on the same trajectory increase monotonically with time, and the range constraint requires that the longitude and latitude of each trajectory point are within the longitude and latitude limit range of the specified area; (4) Conduct a quality assessment of the trajectory subset from the dimension of data integrity, including checking for missing trajectory points and missing values of trajectory points; For any two consecutive trajectory points on a trajectory, if the time interval t between these two points is greater than 2Δt, it is determined that there are missing points in this trajectory, and the number of missing points between these two trajectory points is Δt is the sampling period of this trajectory, that is, the minimum value of the time intervals between all consecutive two points on this trajectory; if any trajectory point has no longitude value, latitude value or timestamp, it is determined that there is a missing value for this trajectory point; (5) Conduct a quality assessment of the trajectory subset from the dimension of data consistency, including checking for trajectory smoothness constraints, trajectory length constraints, abnormal trajectories, abnormal trajectory stop points, and trajectory road network constraints; The trajectory smoothness constraint is to detect the position deviation through a Kalman filter, requiring that there be no sudden position movement of the trajectory points in the trajectory; the trajectory length constraint requires that the cumulative value of the Haversine distances between all consecutive two points in the trajectory ≥ L th , L th is the trajectory length threshold and is set to three times the standard deviation of the average trajectory length; The judgment criterion for the abnormal trajectory is: for any trajectory T in the trajectory subset, if the number of outlier points in trajectory T is greater than the given threshold ρ, then T is determined to be an abnormal trajectory; for any trajectory point p in trajectory T, if the number of neighbor points of p is less than the given threshold η, then p is identified as an outlier point; for any trajectory point q on other trajectories in the trajectory subset, if dist(p,q) ≤ σ, then q is identified as a neighbor point of p, dist(p,q) is the Haversine distance between p and q, and σ is the given distance threshold; The judgment criterion for the abnormal trajectory stay points is as follows: First, all trajectories in the trajectory subset are processed using the TrajDBSCAN algorithm with a small neighborhood to identify all possible stay points. Then, clustering analysis is performed on the identified stay points. The stay points that do not belong to any cluster are determined as abnormal trajectory stay points; The trajectory road network constraint requires that the trajectory must match the path in the road network. The corresponding relationship between each trajectory point in the trajectory and the road section in the road network is found through the OHMM road network matching algorithm. If the corresponding relationship does not exist, it is determined that the corresponding trajectory point and its affiliated trajectory violate the trajectory road network constraint; (6) Conduct quality assessment on the trajectory subset from the dimension of data fairness, including calculating the spatial density of trajectory points in a specific area and the time density of trajectory points in a specific period; For a specific area Q S , the spatial density of the trajectory points within this area where N is the total number of trajectory points in the trajectory subset, N(Q S ) is the number of trajectory points in the trajectory subset that are located within the specific area Q S , Area is the area of the spatial distribution region of the trajectory subset, Area(Q S ) is the area of the specific area Q S ; for a specific time period Q t , the temporal density of the trajectory points within this time period N(Q t ) is the number of trajectory points in the trajectory subset within the specific time period Q t , Time is the overall time span of the trajectory subset, Time(Q t ) is the time span of the specific time period Q t ; (7) According to the statistics in steps (3) to (6), the number and proportion of the detected trajectories and trajectory points, together with the spatial density and time density calculated in step (7), are output as the final quality assessment result.
2. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and characterized in that: The processor is used to execute the computer program to implement a quality assessment method for spatio-temporal trajectory big data as described in claim 1.
3. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements a quality assessment method for spatio-temporal trajectory big data as described in claim 1.
Citation Information
Patent Citations
Data quality evaluation method and device, electronic equipment and storage medium
CN114238267A
Vehicle track editing method and device and terminal equipment
CN118810818A
Vehicle trajectory anomaly detection method based on multi-feature fusion
CN119357866A