A Hierarchical Parallel Clustering Method and System for Massive Trajectories Based on Spatial Similarity

By using spatial similarity, meshing and MinHash signature technology in the trajectory clustering method, the shortcomings of the existing methods in terms of computational complexity and data adaptability are solved, and efficient clustering of massive trajectory data is achieved.

CN119357716BActive Publication Date: 2025-06-17YANTAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411918325.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-06-17
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The existing trajectory clustering methods have shortcomings in terms of computational complexity and data adaptability, especially when processing massive trajectory data, the computational complexity is high and it is difficult to adapt to different data sets.

Method used

The massive trajectory hierarchical parallel clustering method based on spatial similarity is adopted, and the rapid classification and clustering of trajectories are achieved through technologies such as grid unit division, MinHash signature calculation and bucket mapping.

Benefits of technology

This method can adaptively divide trajectories into classes without knowing the number of classes in advance and without training, and improve the clustering efficiency of large-scale trajectory data sets and reduce the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357716B_ABST
    Figure CN119357716B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of spatio-temporal data mining, and particularly to a method and system for hierarchical parallel clustering of massive trajectories based on spatial similarity, including dividing grid cells of a trajectory region according to acquired trajectory data; converting the trajectory data into set data based on the divided grid cells; calculating a MinHash signature corresponding to each trajectory according to the set form of the set data and forming a signature matrix with the MinHash signatures of all trajectories; dividing the obtained signature matrix into several bands and mapping the trajectories in the bands to buckets; dividing the trajectories mapped to the same bucket in at least one band into the same class; without the need to know in advance the number of classes in the trajectory dataset and without training the trajectory dataset, the present invention adaptively and rapidly divides the trajectories into several classes according to their spatial similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spatio-temporal data mining, and particularly to a method and system for hierarchical parallel clustering of massive trajectories based on spatial similarity. Background Art

[0002] Trajectory data usually comes from moving objects (such as animals, people, vehicles, etc.). These objects move in space in chronological order, forming a series of trajectory points. Each trajectory point usually includes spatial position (latitude and longitude) information. The trajectory clustering technology based on spatial similarity is precisely to solve this problem. It divides a set of trajectory data sets into several classes by measuring the spatial similarity between trajectories, so that the trajectories within the same class have a relatively high similarity, while the trajectories between different classes have a relatively low similarity. This method can help us identify trajectories with similar movement patterns and extract valuable information from them. When clustering trajectories based on spatial similarity, it is often necessary to first calculate the similarity between trajectories based on a certain spatial trajectory distance metric method. Then, according to the spatial similarity between trajectories, the trajectories are classified, and the trajectories with higher similarity are divided into one class. Currently, the commonly used trajectory distance metric methods mainly include DTW, Fréchet, Hausdorff, etc. When calculating the distance between a pair of trajectories, these distance metric methods usually take one trajectory as a reference and calculate the Euclidean distance between each trajectory point contained in this trajectory and the trajectory points in the other trajectory. These methods have the following deficiencies:

[0003] (1) The number of trajectory points contained in a trajectory is usually large. Calculating the Euclidean distance between a large number of trajectory point pairs in any two trajectories requires a high time complexity, especially when the number of trajectories is large, which will incur a huge overhead.

[0004] (2) Although some existing work has adopted trajectory compression algorithms, such as the Douglas-Peucker (DP) algorithm, to extract key points from trajectories and represent the trajectories with fewer trajectory points. However, if the number of retained trajectory points is too small, the accuracy of the subsequent clustering results will decrease. On the contrary, it is difficult to solve the problem of high computational complexity between trajectories.

[0005] (3) Since it is impossible to know in advance which class a trajectory belongs to, a large number of unnecessary distance calculations may be required for the trajectories.

[0006] There are also some works on end-to-end clustering models based on deep neural networks to avoid point-by-point distance calculation. However, these models rely heavily on a large amount of labeled training data sets, which may be difficult to obtain quickly in different fields. And these models highly depend on their training data and are difficult to adapt to data sets with different distributions. For example, a model trained on a vehicle trajectory data set in City A may not be well applied to a shared bicycle trajectory data set in City B.

[0007] In summary, most of the existing trajectory clustering methods that do not rely on deep learning require relatively high computational complexity and may be limited by the scale of trajectory data when applied. While the clustering algorithms based on deep learning need to rely on labeled training data sets and are difficult to be flexibly applied to different trajectory data sets when applied. Summary of the Invention

[0008] To solve the above-mentioned problems, the present invention provides a method and system for hierarchical parallel clustering of massive trajectories based on spatial similarity.

[0009] In a first aspect, a method for hierarchical parallel clustering of massive trajectories based on spatial similarity provided by the present invention adopts the following technical solutions:

[0010] A method for hierarchical parallel clustering of massive trajectories based on spatial similarity includes:

[0011] Obtain trajectory data;

[0012] Perform grid cell division of the trajectory area according to the obtained trajectory data;

[0013] Convert the trajectory data into set data based on the divided grid cells;

[0014] According to the set form of the set data, calculate the MinHash signature corresponding to each trajectory and form a signature matrix with the MinHash signatures of all trajectories;

[0015] Divide the obtained signature matrix into several bands and map the trajectories in the bands to buckets;

[0016] Divide the trajectories that are mapped to the same bucket in at least one band into the same class;

[0017] Judge and output the clustering result.

[0018] Further, the performing grid cell division of the trajectory area according to the obtained trajectory data includes, for a given trajectory data set, determining its spatial coverage range based on the minimum and maximum x and y coordinates of all trajectories, and then selecting a fixed grid cell size z to divide the trajectory space into m × n cells.

[0019] Further, the grid cells based on partitioning transform the trajectory data into set data, including determining the grid cells through which the trajectory path passes according to the coordinates of each point in the trajectory. Among them, any trajectory t i is represented as a number of grid cells it passes through, denoted as g i = { c {x,y}}, where x ∈ [1, m] and y ∈ [1, n]. Each grid cell is encoded as the identifier of the grid. Among them, using the Morton coding method, each grid uses an integer as its identifier. Each trajectory corresponds to an integer set composed of the identifiers of several grid cells it passes through.

[0020] Further, calculating the MinHash signature corresponding to each trajectory and forming a signature matrix with the MinHash signatures of all trajectories includes storing the trajectory set data in the form of a matrix to obtain a TE-matrix; calculating the hash value corresponding to each grid cell identifier using a hash function and storing it in the form of a matrix to obtain an HV-matrix; generating a Sig-matrix based on the TE-matrix and the HV-matrix, and initializing the elements in the Sig-matrix to ∞. Among them, traverse each row i in the TE-matrix, check the value of each column. If the value is '0', do nothing; if the value is '1', update the corresponding column in the Sig-matrix with the hash value in the same row of the TE-matrix; and update the Sig-matrix with the rows of the HV-matrix. Finally, the signature matrix is obtained.

[0021] Further, dividing the obtained signature matrix into several bands and mapping the trajectories in the bands to buckets includes further dividing the rows of the Sig-matrix into multiple bands. For each band, create a bucket array, and map the trajectories with the same signature in the band to the same bucket. Trajectories in the same bucket indicate that there is a high Jaccard similarity between their signatures.

[0022] Further, dividing the trajectories that are mapped to the same bucket in at least one band into the same class includes creating an initial class C 1 and setting it to any bucket in the first band. Scan the buckets in the remaining bands to find buckets that intersect with the trajectories in C 1 . If there are such buckets, merge the trajectories in the buckets into C 1 . After scanning each bucket in the second to the last band, return to the first band. At this time, if buckets that intersect with C 1 are found, merge them as well. After scanning the buckets in the first band, it can be confirmed C1 After determining C1, the unprocessed buckets in the first band are sequentially used as new classes until all the buckets in the first band are processed. Then, it is determined whether there is an intersection between each of the finally obtained classes. If there is an intersection, the classes with intersections are also merged. The classes that finally have no intersections with each other are the final result of this clustering process.

[0023] Furthermore, the determination and output of the clustering result include determining whether each class meets the set conditions. If it meets the set conditions, the final clustering result is output; if not, the clustering operation is repeated for each class in parallel, and when dividing the grid cells, a smaller grid cell is used to divide the entire trajectory space of the trajectories in this class.

[0024] In a second aspect, a hierarchical parallel clustering system for massive trajectories based on spatial similarity includes:

[0025] A data acquisition module, configured to acquire trajectory data;

[0026] A unit division module, configured to divide the grid cells of the trajectory area according to the acquired trajectory data;

[0027] A conversion module, configured to convert the trajectory data into set data based on the divided grid cells;

[0028] A matrix module, configured to calculate the MinHash signature corresponding to each trajectory according to the set form of the set data and form a signature matrix with the MinHash signatures of all trajectories;

[0029] A mapping module, configured to divide the obtained signature matrix into several bands and map the trajectories in the bands to buckets;

[0030] A division module, configured to divide the trajectories that are mapped to the same bucket in at least one band into the same class;

[0031] A judgment and output module, configured to judge and output the clustering result.

[0032] In a third aspect, the present invention provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions are adapted to be loaded and executed by a processor of a terminal device for the hierarchical parallel clustering method for massive trajectories based on spatial similarity.

[0033] In a fourth aspect, the present invention provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement each instruction; the computer-readable storage medium is used to store multiple instructions, and the instructions are adapted to be loaded and executed by the processor for the hierarchical parallel clustering method for massive trajectories based on spatial similarity.

[0034] In summary, the present invention has the following beneficial technical effects:

[0035] (1) Without the need to know in advance the number of classes in the trajectory dataset and without training the trajectory dataset, the present invention adaptively and quickly divides the trajectories into several classes according to their spatial similarity. And after the division, the spatial similarity between the trajectories within each class is relatively high, while the spatial similarity between the trajectories in different classes is relatively low.

[0036] (2) In order to improve the clustering efficiency of large-scale trajectory datasets, a method of converting trajectory data into set data and clustering according to the local sensitive hashing method is proposed, effectively avoiding the high time complexity problem caused by the distance calculation between trajectory pairs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Schematic diagram of a method for hierarchical parallel clustering of massive trajectories based on spatial similarity according to Embodiment 1 of the present invention;

[0038] Figure 1 In (a) is a flowchart of a method for hierarchical parallel clustering of massive trajectories based on spatial similarity according to Embodiment 1 of the present invention;

[0039] Figure 1 In (b) is a flowchart of a classification method based on bucket mapping according to Embodiment 1 of the present invention;

[0040] Figure 2 Example diagram of grid cell division of the trajectory space according to Embodiment 1 of the present invention;

[0041] Figure 3 Example diagram of encoding grid cells according to Embodiment 1 of the present invention;

[0042] Figure 4 According to Embodiment 1 of the present invention, after grid cell division and encoding, Figure 2 Example diagram of converting trajectory data into set data;

[0043] Figure 5 Schematic diagram of the calculation process of the TE-matrix and HV-matrix according to Embodiment 1 of the present invention;

[0044] Figure 5 In (a) is a schematic diagram of the trajectory encoding matrix (TE-matrix) according to Embodiment 1 of the present invention;

[0045] Figure 5 In (b) is the matrix (HV-matrix) obtained from the trajectory encoding according to the hash function according to Embodiment 1 of the present invention;

[0046] Figure 5Among them, (c) is the hash function and trajectory encoding of Embodiment 1 of the present invention;

[0047] Figure 6 It is a schematic diagram of the calculation example of the MinHash signature matrix in Embodiment 1 of the present invention;

[0048] Figure 7 It is a schematic diagram of clustering based on the MinHash signature matrix in Embodiment 1 of the present invention;

[0049] Figure 7 Among them, (a) is a schematic diagram of dividing the obtained MinHash signature matrix into several bands in Embodiment 1 of the present invention;

[0050] Figure 7 Among them, (b) is a schematic diagram of mapping the trajectories in the band to the buckets in Embodiment 1 of the present invention;

[0051] Figure 7 Among them, (c) is a schematic diagram of classifying trajectories based on the bucket mapping result in Embodiment 1 of the present invention;

[0052] Figure 8 It is a schematic diagram of the influence of the number of hash functions in Embodiment 1 of the present invention on the performance of the HPTC algorithm;

[0053] Figure 8 Among them, (a) shows the influence of the number of hash functions in Embodiment 1 of the present invention on the calculation time when the HPTC algorithm performs clustering on four trajectory datasets;

[0054] Figure 8 Among them, (b) shows the influence of the number of hash functions in Embodiment 1 of the present invention on the accuracy (referring to the DBI index) when the HPTC algorithm performs clustering on four trajectory datasets;

[0055] Figure 8 Among them, (c) shows the influence of the number of hash functions in Embodiment 1 of the present invention on the accuracy (referring to the SC index) when the HPTC algorithm performs clustering on four trajectory datasets;

[0056] Figure 9 It is a schematic diagram of the influence of the grid cell size in Embodiment 1 of the present invention on the performance of the HPTC algorithm;

[0057] Figure 9 Among them, (a) shows the influence of the grid cell size in Embodiment 1 of the present invention on the calculation time and accuracy (referring to the DBI index) when the HPTC algorithm performs clustering on the LABOMNI dataset;

[0058] Figure 9Among them, (b) shows the influence of the grid cell size in Embodiment 1 of the present invention on the calculation time and accuracy of the HPTC algorithm when clustering on the LABOMNI dataset (with the DBI index as a reference);

[0059] Figure 10 is a schematic diagram showing the influence of the DP algorithm in Embodiment 1 of the present invention on the performance of the HPTC algorithm;

[0060] Figure 10 Among them, (a) shows the influence of the DP algorithm in Embodiment 1 of the present invention on the accuracy of the HPTC algorithm when clustering on four trajectory datasets (with the SC index as a reference);

[0061] Figure 10 Among them, (b) shows the influence of the DP algorithm in Embodiment 1 of the present invention on the accuracy of the HPTC algorithm when clustering on four trajectory datasets (with the DBI index as a reference);

[0062] Figure 10 Among them, (c) shows the influence of the DP algorithm in Embodiment 1 of the present invention on the calculation time of the HPTC algorithm when clustering on four trajectory datasets;

[0063] Figure 11 is a schematic diagram showing the performance comparison between the HPTC algorithm and the benchmark algorithm in Embodiment 1 of the present invention;

[0064] Figure 11 Among them, (a) shows the comparison of the calculation time when the HPTC algorithm in Embodiment 1 of the present invention clusters on four trajectory datasets with three benchmark methods;

[0065] Figure 11 Among them, (b) shows the comparison of the accuracy of the results (with the DBI index as a reference) when the HPTC algorithm in Embodiment 1 of the present invention clusters on four trajectory datasets with three benchmark methods;

[0066] Figure 11 Among them, (c) shows the comparison of the accuracy of the results (with the SC index as a reference) when the HPTC algorithm in Embodiment 1 of the present invention clusters on four trajectory datasets with three benchmark methods. Detailed implementation manners

[0067] The present invention will be further described in detail below with reference to the accompanying drawings.

[0068] Embodiment 1

[0069] Term explanation:

[0070] (1) Trajectory: A trajectory t i is defined as a finite sequence of points with x and y coordinates, denoted as ti = {P1, P2,... , P z}, where P j represents the t i th j point of the trajectory.

[0071] (2) Trajectory clustering: Given a set of trajectories T = { t 1 , t 2 , ... , t n}, trajectory clustering is to divide T into k clusters C 1 , C 2 , ... , C k based on a specific similarity measure, satisfying the following conditions: Each cluster C i is a subset of the trajectories in T, i.e., C i ⊆ T; (2) Each trajectory t i belongs to only one cluster, satisfying and = (when i ≠ j ); (3) Trajectories within the same cluster C i have high similarity, while trajectories in different clusters have low similarity.

[0072] As Figure 1 shown, this embodiment discloses a method for hierarchical parallel clustering of massive trajectories based on spatial similarity, including:

[0073] Step 1: Sequentially read trajectory data from the trajectory database;

[0074] Step 2: Determine the trajectory space where the trajectory data in the trajectory dataset is located, and then divide the entire space into grid cells of the same size;

[0075] Specifically, for a given trajectory dataset, we first determine its spatial coverage based on the minimum and maximum x and y coordinates of all trajectories. Then, select a fixed grid cell size z and divide the trajectory space into m × n cells. Figure 2 shows an example of using grid division for the trajectory space.

[0076] Step 3: Convert the trajectory data into set data based on the identifiers of the sequence of grid cells passed by the trajectory;

[0077] Step 301: The grid cells passed by the trajectory path can be determined according to the coordinates of each point in the trajectory. Any trajectory t i is represented as a number of grid cells it passes through, denoted as g i = { c {x,y}}, where x ∈ [1, m] and y ∈ [1, n]. Figure 4 The second column shows a specific example of re-representing the trajectory data in Figure 2 as the sequence of grid cells it passes through.

[0078] Step 302: For simplicity of representation, each grid cell is encoded as the identifier of this grid. Specifically, Figure 3 shows an example of Morton encoding. Using the encoding method of Morton encoding, each grid uses an integer as its identifier.

[0079] Step 303: Each trajectory corresponds to a set of integers composed of the identifiers of several grid cells it passes through, Figure 4 The third column shows a specific example of converting the trajectory data in Figure 2 into set data.

[0080] Step 4: According to the trajectories in set form, calculate the MinHash signature corresponding to each trajectory based on a certain number of hash functions and form a signature matrix with the MinHash signatures of all trajectories;

[0081] The generation process of the MinHash signature matrix includes:

[0082] Step 401: Store the trajectory set in the form of a matrix, and the obtained matrix is called the "TE-matrix".

[0083] Figure 5 The (a) in Figure 4 shows the "TE-matrix" of the trajectories in t i j (i ∈ [1, ) contains the grid cell identifier c j ( c j ∈ I), then the corresponding element ei,j Set it to 1, otherwise 0. The "TE-matrix" describes the relationship between the trajectory and the identifiers it contains, thus simplifying the process of calculating the signature using MinHash.

[0084] Step 402: Calculate the hash values corresponding to each grid cell identifier using a certain number of hash functions and store them in the form of a matrix, called the "HV-matrix". Figure 5 Figure (b) in [reference] shows an example of an "HV-matrix", introducing 4 hash functions, as shown in Figure 5 Figure (c) in [reference]. Calculate the hash values corresponding to each hash function for each cell identifier in turn and store them in the "HV-matrix". The general format of these hash functions is: . Where, a , b and c are random variables, j represents the encoding of the grid cell c j ( c j ∈I), 2 c -1 represents a series of prime numbers.

[0085] Step 403: Generate a signature matrix for the "TE-matrix", called the "Sig-matrix". The columns of the "Sig-matrix" correspond to the trajectories, and the rows correspond to L hash functions h 1 (x) , h 2 (x) , ..., h L (x) .

[0086] Figure 6 Figure [reference] shows the process of generating the "Sig-matrix" based on the "TE-matrix" and the "HV-matrix". For simplicity of explanation, let Figure 5 ([[]] sig (j,z) ( j ∈[1, l], z ∈[1, ) represents the element of the z-th trajectory corresponding to the j -th hash function in the "Sig-matrix". The specific generation process of the "Sig-matrix" is as follows:

[0087] Step 4031: Initialize the elements in the "Sig-matrix" to ∞.

[0088] Step 4032: Traverse each row i in the "TE-matrix" and check the values in each column. If the value is '0', no operation is performed; if the value is '1', update the corresponding column in the "Sig-matrix" with the hash value in the same row of the "TE-matrix". Since the number of columns in the "HV-matrix" is equal to the number of rows in the "Sig-matrix", the rows of the "HV-matrix" can be directly used to update the "Sig-matrix". Specifically, if the hash value h j (r) is less than the current value sig (j,z) , then sig (j,z) is updated to h j (r) .

[0089] Step 5: Divide the obtained signature matrix into several bands, and for each band, use an independent bucket array to map the trajectories with the same signature in each band to the same bucket.

[0090] As shown in (a) and (b) of Figure 7 , to group the trajectories based on the determined "Sig-matrix", the following steps are performed:

[0091] Step 501: As shown in (a) of Figure 7 , further divide the rows of the "Sig-matrix" into multiple bands.

[0092] Step 502: As shown in (b) of Figure 7 , for each band, create a bucket array and map the trajectories with the same signature in the band to the same bucket. Trajectories in the same bucket indicate that there may be a high Jaccard similarity between their signatures, suggesting that they may be strongly similar in space.

[0093] Step 6: Divide the trajectories that are mapped to the same bucket in at least one band into the same class;

[0094] Figure 1 (b) of Figure 7 shows a detailed flowchart of this step,

[0095] (c) of C 1 shows a specific example, based on the following steps:

[0096] Step 601: Create an initial class C 1Buckets with intersecting trajectories within it, if there are such buckets, then merge the trajectories in the buckets into C 1 .

[0097] Step 603: Although the buckets in the same strip do not share trajectories, C 1 have been merged with some buckets in other strips, and thus may intersect with the unprocessed buckets in the first strip. After scanning each bucket in the second to the last strip, return to the first strip. At this time, if it is found that there are buckets intersecting with C 1 , merge them as well. After scanning the buckets in the first strip, it can be confirmed C 1 .

[0098] Step 604: After determining C 1 , sequentially take the unprocessed buckets in the first strip as new classes, and repeat the process of Steps 602 - 603 until all the buckets in the first strip are processed.

[0099] Step 605: Determine whether there is an intersection between each of the finally obtained classes. If there is, merge the classes with intersections as well. The finally obtained classes that do not intersect with each other are the final results of this clustering process.

[0100] Step 7: Determine whether each class meets the conditions. If it meets the conditions, output the final clustering result; if not, repeat Steps 1 - 6 for each class in parallel, where when performing Step 2, divide the entire trajectory space of the trajectories in this class using smaller grid cells.

[0101] The specific condition here means that after further clustering, the trajectories in this class are not divided into more classes. To balance accuracy and efficiency, HPTC initially converts the trajectory data into a set of several grid identifiers based on larger grid cells, so as to represent the trajectories using fewer identifiers, which can reduce the computational amount when calculating the similarity between trajectories subsequently. This process will generate a coarse - grained clustering result, which may lead to relatively low similarity within some of these classes and requires further refinement. To achieve more refined clustering, the system re - divides the trajectory space in these classes using grid cells that are half the size of the original grid cells, further refines the trajectory representation, and increases the length of each trajectory representation to capture more spatial features. Then, further clustering is performed on this class, thus generating more sub - clusters with high cohesion. If further refinement is needed, this process will continue. If a cluster cannot be divided into more cohesive sub - clusters, it is considered that the cluster has achieved sufficient intra - cluster similarity and no further clustering is required.

[0102] The effectiveness of the present invention is demonstrated by experiments below.

[0103] The experiments of the present invention were conducted on a computer equipped with an Intel Core i5 2.5 GHz processor, a 64-bit Windows 10 operating system, and 8 GB of RAM. All algorithms were implemented in Java.

[0104] 1. Experimental settings and datasets.

[0105] Datasets. Four trajectory datasets were mainly used in this experiment, namely: LABOMNI, T-Drive, Geolife, and NGSIM. Specifically, LABOMNI contains 209 trajectories and provides standardized clustering results, forming a total of 15 standard clusters. Geolife includes data from 182 users, with a total of 17,621 trajectories, covering approximately 1.2 million kilometers and a total duration of over 48,000 hours. T-Drive contains the trajectories of 10,357 taxis in Beijing, with a total of 15 million GPS coordinates, covering over 9 million kilometers. After data cleaning, the final dataset contains 207,966 trajectories. NGSIM contains 9,525 trajectories from Highway 101 in the United States, from 7:50 to 8:05 in the morning.

[0106] Benchmark algorithms. Improved DBSCAN, K -Swaps, and LSH-ED were used as benchmark algorithms, and these four benchmark algorithms will be briefly introduced below.

[0107] The Improved DBSCAN algorithm is a compression-based and density-based clustering method based on Douglas-Peucker (DP). This method mainly consists of two main steps: trajectory compression and grouping. To improve the efficiency and accuracy of DTW similarity measurement, the DP algorithm is used to compress the trajectories, and then the similarity between the compressed trajectories is calculated. In the grouping stage, the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm is used and statistical methods are used to determine the DBSCAN parameters.

[0108] The K-Swaps algorithm first applies K -Set algorithm to obtain an initial solution, and performs swaps by replacing the histogram of a randomly selected cluster with a random data object. Then, the K -Set algorithm is applied again, and the objective function SDH (sum of the minimum distances to the histogram) is calculated. If the new SDH is smaller than the previous result, the new solution is accepted and used in the next iteration. This process is repeated for a fixed number of iterations, which is the only parameter given by the user. Among them KThe -Set algorithm consists of two steps: assignment and update. First, randomly select K data objects as the initial representatives. At this stage, the representatives are just the histograms of the selected data objects. In the assignment step, each data object is assigned to its nearest histogram. In the design of our algorithm, the trajectory data can be converted into a set form in the preprocessing stage. Therefore, we can use the trajectory data converted into a set form K -Swaps algorithm for clustering calculation.

[0109] The LSH-ED algorithm proposes a clustering based on LSH in Euclidean space. The condition for applying LSH-ED to trajectory data is that all trajectories must be of the same length. To overcome this problem, the DP algorithm is improved in the paper to compress the trajectories to a fixed length and minimize the loss as much as possible. Then the trajectories are mapped into different buckets by using a set of hash functions.

[0110] 2. HPTC performance evaluation.

[0111] In this embodiment, the influence of different parameters on the clustering efficiency and accuracy of HPTC, such as the number of hash functions L, the grid cell size Δd, and the DP algorithm, is analyzed for HPTC; then the clustering efficiency and accuracy are compared with those of the baseline algorithm.

[0112] 1) Evaluation metrics for trajectory clustering.

[0113] The silhouette coefficient (SC) is an evaluation metric used to evaluate the distance between a target trajectory and other trajectories within its cluster and the clustering with the nearest trajectory in other clusters. The result is presented as a ratio within the range of [-1, 1]. A higher SC value indicates a better clustering effect.

[0114] The Davies-Bouldin index (DBI) is another evaluation metric that also uses the ratio of the distance between clusters and the distance within clusters. A lower DBI value indicates a better clustering effect.

[0115] 2) Clustering efficiency of the HPTC algorithm under different parameters.

[0116] (1) We first evaluate the clustering accuracy and efficiency of HPTC under different numbers of hash functions L. Figure 8 in (b) and Figure 8 in (c) respectively show the effects of the clustering accuracy measured by the DBI and SC metrics on all datasets. The higher the SC value, the better the clustering accuracy, while the lower the DBI value, the more superior the clustering performance. As Figure 8As shown in (c), as L increases, the SC value gradually rises, indicating that the clustering performance of HPTC is improved by increasing the number of hash functions. However, when L exceeds a certain threshold (e.g., L = 75 in the T-Drive dataset), the SC value tends to stabilize and does not show a significant increase. This indicates that even if more hash functions are added, the clustering performance will not be further improved. The reason for this phenomenon is that when L is small, the number of rows in each band is small, which may cause dissimilar trajectories to be mapped to the same bucket, thus reducing the clustering accuracy. Increasing the value of L at this stage will increase the number of rows in each band, thereby reducing the possibility of dissimilar trajectories being assigned to the same cluster and improving the accuracy. However, if L is too large, it may cause similar trajectories to be scattered into different buckets, thus weakening the distinguishability between clusters and ultimately resulting in no significant improvement in clustering accuracy. For the same reason, as L increases, the clustering accuracy measured by DBI also shows a similar trend on each dataset, as shown in Figure 8 shown in (b).

[0117] Figure 8 As shown in (a), as the number of hash functions L increases, the processing time increases significantly. This is because using more hash functions will undoubtedly increase the number of hash values to be processed, resulting in a greater computational overhead. In this test, it was found that the trends of clustering accuracy and efficiency are similar on different datasets. Therefore, the following evaluations will focus on the LABOMNI and NGSIM datasets, which are the smallest and largest test datasets respectively.

[0118] (2) The impact of the cell size (Δd) on clustering accuracy and time cost was evaluated. The cell size depends on the spatial coverage of the dataset where the trajectories are located. Specifically, the spatial coverage of each dataset is defined as a rectangular area, where the length of the rectangle is obtained by calculating the difference between the maximum x-coordinate and the minimum x-coordinate among all trajectory points, and the width is determined by the same calculation method for the y-coordinates. According to this method, the rectangular space of the LABOMNI dataset is approximately 150×150, while the spatial coverage of the NGSIM dataset is approximately 2700×2400. Therefore, we chose a smaller cell size for LABOMNI and a relatively larger cell size for NGSIM.

[0119] As Figure 9As shown, we initially set the cell sizes of LABOMNI and NGSIM to 35 and 100 respectively. Subsequently, we halved the cell size in subsequent steps and measured the DBI values. To better observe the impact of cell size on the performance of the HPTC algorithm, no iteration was performed during the clustering process. In both datasets, as the cell size decreased, the DBI value decreased significantly, indicating an improvement in clustering performance. This supports our conclusion that smaller cells enhance the Jaccard similarity between trajectory encodings, thereby improving the similarity of trajectories within each cluster. However, when the cell size (Δd) was further reduced and exceeded a certain threshold, the DBI value tended to stabilize, indicating that further reducing the cell size had limited improvement on clustering performance. In addition, continuously reducing the cell size would significantly increase the clustering cost. Therefore, the cell size should be appropriately selected to balance clustering accuracy and computational efficiency.

[0120] (3) The performance of the HPTC algorithm was tested under two conditions: one was using the original unsimplified trajectories, and the other was using the trajectories simplified by the DP algorithm for clustering. As Figure 10 shown, trajectory clustering after extracting key points did not reduce the accuracy; in fact, in some datasets (such as LABOMNI and Geolife), the clustering performance was improved, as shown in Figure 10 (a) in Figure 10 and Figure 10 (b) in

[0121] 3. Comparison with benchmark algorithms.

[0122] Figure 11 (a) in Figure 11 shows the clustering time cost of the HPTC algorithm and the benchmark methods on each dataset. Obviously, in all datasets, the HPTC algorithm is superior to the Improved DBSCAN and K-Swaps algorithms in terms of running efficiency. In larger datasets (such as Geolife and T-Drive), the running time of HPTC is significantly reduced compared to the Improved DBSCAN and K-Swaps algorithms, reducing by approximately two orders of magnitude. This is because HPTC converts trajectories into integer sets and does not need to spend a lot of time calculating pairwise distances between trajectories in similarity evaluation. Although the running time of HPTC is slightly longer than that of the LSH-ED algorithm, as shown in Figure 11 (b) and Figure 11As shown in (c), the clustering accuracy of HPTC on all datasets is significantly better than that of LSH-ED.

[0123] First, the DBI metric is used to evaluate the clustering accuracy of different algorithms on each dataset. A lower DBI value indicates better clustering performance. The results are as Figure 11 shown in (b). It can be seen that the HPTC method achieves the minimum DBI value on the LABOMNI dataset, and the clustering results are highly consistent with the true labels. On the T-Drive, Geolife, and NGSIM datasets, the performance of HPTC is significantly better than that of LSH-ED and K-Swaps, and it is only slightly inferior to Improved DBSCAN. However, in terms of clustering cost, as Figure 11 shown in (b), the performance of HPTC on the T-Drive, Geolife, and NGSIM datasets is much better than that of Improved DBSCAN, and the efficiency is improved by nearly two orders of magnitude. Therefore, our HPTC algorithm shows more superior overall performance. Especially when dealing with large-scale datasets, HPTC can significantly accelerate the clustering speed while maintaining high clustering accuracy.

[0124] Then, the SC metric is used as the evaluation criterion for clustering accuracy to compare the clustering accuracy. The results are as Figure 11 shown in (c), where a larger SC value indicates higher clustering accuracy. The conclusion is basically the same as that obtained using DBI as the evaluation metric. Our algorithm performs better than all benchmark algorithms on the Geolife dataset and is only slightly inferior to the Improved DBSCAN algorithm on the other three datasets.

[0125] Example 2

[0126] This example provides a hierarchical parallel clustering system for massive trajectories based on spatial similarity, including:

[0127] A data acquisition module configured to acquire trajectory data;

[0128] A unit division module configured to perform grid cell division of the trajectory area according to the acquired trajectory data;

[0129] A conversion module configured to convert the trajectory data into set data based on the divided grid cells;

[0130] A matrix module configured to calculate the MinHash signature corresponding to each trajectory according to the set form of the set data and form a signature matrix with the MinHash signatures of all trajectories;

[0131] A mapping module configured to divide the obtained signature matrix into several bands and map the trajectories in the bands to buckets;

[0132] A partitioning module, configured to partition trajectories mapped to the same bucket in at least one band into the same class;

[0133] A judgment output module, configured to judge whether each class meets the conditions and output a clustering result.

[0134] A computer-readable storage medium storing multiple instructions, the instructions being adapted to be loaded and executed by a processor of a terminal device for the described method for hierarchical parallel clustering of massive trajectories based on spatial similarity.

[0135] A terminal device, comprising a processor and a computer-readable storage medium, the processor being used to implement each instruction; the computer-readable storage medium being used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor for the described method for hierarchical parallel clustering of massive trajectories based on spatial similarity.

[0136] The above are all preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention should be covered within the protection scope of the present invention.

Claims

1. A hierarchical parallel clustering method for massive trajectories based on spatial similarity, characterized in that: include: Get trajectory data; Divide the track area into grid cells according to the acquired track data; Convert trajectory data into set data based on divided grid cells; According to the collection form of the collection data, the MinHash signature corresponding to each trajectory is calculated and the MinHash signatures of all trajectories are combined into a signature matrix; The obtained signature matrix is ​​divided into several bands, and the trajectories in the bands are mapped into buckets; Trajectories that are mapped to the same bucket in at least one band are grouped into the same class; Determine and output clustering results; The grid unit division of the trajectory area according to the acquired trajectory data includes, for a given trajectory data set, determining its spatial coverage based on the minimum and maximum x and y coordinates of all trajectories, and then selecting a fixed grid unit size z to divide the trajectory space into m × n units; The method converts the trajectory data into set data based on the divided grid units, including determining the grid units that the trajectory path passes through according to the coordinates of each point in the trajectory, wherein any trajectory t i is represented by the number of grid cells it passes through, denoted by g i ={ c {x,y} }, where x∈[1, m] and y∈[1, n], encode each grid cell as the identifier of the grid, where each grid is encoded using an integer as its identifier using the Morton encoding method, and each trajectory corresponds to an integer set consisting of the identifiers of several grid cells it passes through; The method comprises calculating the MinHash signature corresponding to each trajectory and forming a signature matrix with the MinHash signatures of all trajectories, including storing the trajectory set data in the form of a matrix to obtain a TE-matrix; using a hash function to calculate the hash value corresponding to each grid unit identifier and storing it in the form of a matrix to obtain an HV-matrix; generating a Sig-matrix based on the TE-matrix and the HV-matrix, initializing the elements in the Sig-matrix to ∞, wherein each row i in the TE-matrix is ​​traversed, and the value of each column is checked. If the value is '0', no operation is performed; if the value is '1', the hash value of the same row in the TE-matrix is ​​used to update the corresponding column in the Sig-matrix; and the rows of the HV-matrix are used to update the Sig-matrix to finally obtain the signature matrix; The obtained signature matrix is ​​divided into a plurality of bands, and the trajectories in the bands are mapped into buckets, including further dividing the rows of the Sig-matrix into a plurality of bands, creating a bucket array for each band, and mapping the trajectories with the same signature in the band to the same bucket, and the trajectories in the same bucket indicate that the signatures have high Jaccard similarity; The method of classifying the tracks mapped to the same bucket in at least one band into the same class includes creating an initial class C 1 And set it to any bucket in the first band, scan the buckets in the remaining bands, and find the bucket that matches C 1 There are buckets where the trajectories in the bucket intersect. If there is such a bucket, merge the trajectories in the bucket into C 1 After scanning each bucket in the second to last band, return to the first band. C 1 If there are intersecting buckets, merge them as well. After scanning the buckets in the first band, you can confirm C 1 After determining C1, the unprocessed buckets in the first band are taken as new classes in turn until all the buckets in the first band are processed. It is determined whether there is an intersection between each of the final classes. If there is, the classes with intersections are also merged. The classes with no intersections are the final results of this clustering process. The judging and outputting of clustering results includes judging whether each class satisfies a set condition, and if so, outputting the class as a class in the final clustering result; if not, repeating the clustering operation for each class in parallel, wherein a smaller grid unit is used to divide the entire trajectory space of the trajectory in the class when performing grid unit division; the condition refers to that after further clustering, the trajectories in the class are not divided into more subclasses, specifically, the trajectories in each class are grid-divided using a grid unit of half the size of the original grid unit to obtain a finer-grained trajectory representation, and then further clustering is performed in parallel within each class at a finer granularity; if a class is not further divided into multiple subclasses during the optimization process, it is considered that the class has achieved sufficient intra-class similarity and is output as the final class; otherwise, a finer-grained clustering optimization process is continued to be recursively performed on these subclasses until further division is impossible.

2. A massive trajectory hierarchical parallel clustering system based on spatial similarity, executing the method as claimed in claim 1, characterized in that: include: A data acquisition module is configured to acquire trajectory data; A unit division module is configured to divide the track area into grid units according to the acquired track data; The conversion module is configured to convert the trajectory data into set data based on the divided grid units; The matrix module is configured to, according to the set form of the set data, calculate the MinHash signature corresponding to each trajectory and form the MinHash signatures of all trajectories into a signature matrix; A mapping module, configured to divide the obtained signature matrix into a number of bands and map the trajectories in the bands into buckets; A partitioning module configured to partition trajectories mapped to the same bucket in at least one band into the same class; The judgment output module is configured to judge and output the clustering result.

3. A computer-readable storage medium storing a plurality of instructions, characterized in that: The instructions are suitable for being loaded by a processor of a terminal device and executing the method according to claim 1 .

4. A terminal device, comprising a processor and a computer-readable storage medium, wherein the processor is used to implement each instruction; and the computer-readable storage medium is used to store multiple instructions, characterized in that: The instructions are suitable for being loaded by a processor and executing the method as claimed in claim 1 .

Citation Information

Patent Citations

  • Track clustering method

    CN116881750A