Parallel online similarity analysis method for trajectory flow

Through the parallel online similarity analysis method, combined with load partitioning, dynamic load balancing and multi-level pruning technology, the problems of outdated results and high calculation costs in the existing online trajectory similarity connection methods are solved, real-time dynamic updates and efficient calculations are achieved.

CN120145071APending Publication Date: 2025-06-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510273160.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing online trajectory similarity connection methods cannot be updated dynamically in real time, resulting in a large number of outdated trajectory pairs in the connection results, and the computational cost is too high when processing large-scale data sets to support real-time updates or scaling.

Method used

A parallel online similarity analysis method for trajectory flow is proposed, which realizes incremental similarity update through load partitioning and dynamic load balancing, and introduces time-aware exponential decay factor to eliminate outdated results. In addition, multi-level pruning technology is used to accelerate similarity calculations in both spatial and temporal dimensions.

Benefits of technology

Real-time dynamic update of trajectory connection results is realized, eliminating outdated results, improving computing efficiency, reducing data redundancy, and supporting real-time updates of large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145071A_ABST
    Figure CN120145071A_ABST
Patent Text Reader

Abstract

The invention discloses a parallel online similarity analysis method for trajectory flows, and relates to the technical field of trajectory analysis, and the method comprises the following steps: S1, collecting a first trajectory position flow and a second trajectory position flow of a moving object, and carrying out load partitioning to obtain a plurality of partitioning matrixes; s2, carrying out load balancing processing on each partition matrix; s3, based on the partition matrix subjected to load balancing processing, similarity increment updating is carried out; and S4, on the basis of a similarity increment updating result, calculating trajectory similarity by using an attenuation factor when new data is added. In order to improve the efficiency, a multi-stage pruning technology is implemented in two dimensions of space and time, and similarity calculation is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of trajectory analysis, and particularly relates to a parallel online similarity analysis method for trajectory streams. Background Art

[0002] Due to the increasing popularity of devices with GPS capabilities and location-based services, the amount of trajectory data has been growing rapidly. Efficiently processing and analyzing large-scale trajectory data through basic operations such as trajectory similarity search and similarity join has become an important research direction in the field of data science. These efforts lay the foundation for various location-based services, such as anomaly detection, carpooling recommendations, and location-based data cleaning.

[0003] In recent years, the maturity and widespread application of real-time data processing frameworks have led to a significant change in the use of trajectories. The current mainstream model emphasizes the real-time collection, processing, and feedback of data. Researchers no longer only focus on how to effectively perform offline batch processing of trajectory data, but instead pay more attention to achieving low-latency and high-throughput online computing by leveraging parallel computing resources. Some existing research has developed in-memory online trajectory similarity query and similarity join algorithms.

[0004] In existing trajectory similarity joins, the similarity between two trajectories is static, and when their location information is not updated, the similarity does not change. All trajectory pairs whose similarity once exceeded a given threshold are stored in the join result. As the processing clock (i.e., the current system timestamp) advances, the join results calculated by these solutions inevitably contain a large number of outdated trajectory pairs, making them inapplicable to dynamically evolving trajectories. To eliminate outdated results and maintain real-time trajectory joins, online trajectory similarity joins have been proposed.

[0005] Existing research related to online trajectory search uses a sliding window to define a predefined time range in the trajectory stream and perform similarity joins within this window, thus ignoring the trajectory points outside the window. The query efficiency highly depends on the window size. However, in practical applications, it is not easy to provide an accurate window length in advance. For example, in the case of long-term observations such as detecting animal migration patterns or tropical cyclone paths, if the window size is too large, the scalability of these methods is limited and the computational cost is too high. Therefore, they cannot support real-time updates or scale to large-scale datasets. In addition, these methods do not consider the decay of outdated trajectory pairs. Therefore, their solutions cannot be directly used to solve problems in an exponentially decaying setting. Summary of the Invention

[0006] To solve the above problems, the present invention proposes a parallel online similarity analysis method for trajectory streams.

[0007] The technical solution of the present invention is: A parallel online similarity analysis method for trajectory streams includes the following steps:

[0008] S1. Collect the first trajectory position stream and the second trajectory position stream of moving objects, and perform load partitioning to obtain a number of partition matrices;

[0009] S2. Perform load balancing processing on each partition matrix;

[0010] S3. Based on the partition matrices after load balancing processing, perform similarity incremental update;

[0011] S4. Based on the similarity incremental update result, calculate the trajectory similarity using a decay factor when new data is added.

[0012] Further, in S1, the positions of the first trajectory position stream are partitioned in a column-cyclic manner, and the positions of the second trajectory position stream are partitioned in a row-cyclic manner to obtain a number of partition matrices of size , where represents the number of available threads for downstream trajectory similarity connection.

[0013] Further, S2 includes the following sub-steps:

[0014] S21. Construct a load matrix with the same size as the partition matrix;

[0015] S22. Calculate the first imbalance factor and the second imbalance factor according to the load matrix;

[0016] S23. Perform load balancing processing on the partition matrix according to the first imbalance factor and the second imbalance factor.

[0017] Further, in S22, the calculation formula of the first imbalance factor is:

[0018] ;

[0019] ;

[0020] ;

[0021] ;

[0022] In the formula, represents the total partition load of the th column in the load matrix, represents the total partition load of the th row in the load matrix, represents the maximum value operation, represents the minimum value operation, Indicates the number of available threads for downstream trajectory similarity joining, Indicates the row and column load of the load matrix, Indicates the row and column load of the load matrix, Indicates the dimension of the load matrix;

[0023] In S22, the calculation formula for the second imbalance factor is as follows:

[0024] ;

[0025] In the formula, Indicates the sum of the loads of the row of the load matrix, Indicates the sum of the loads of the row of the load matrix.

[0026] Furthermore, in S23, if the first imbalance factor is greater than or equal to the threshold, the subsequent positions of the first trajectory position stream are assigned to all partition matrices corresponding to the column with the minimum load until the first imbalance factor is less than the threshold;

[0027] In S23, if the second imbalance factor is greater than or equal to the threshold, the subsequent positions of the first trajectory position stream are assigned to all partition matrices corresponding to the row with the minimum load until the second imbalance factor is less than the threshold.

[0028] Furthermore, S3 includes the following sub-steps:

[0029] S31. When adding a new position, perform global pruning on the trajectory pairs;

[0030] S32. Perform local pruning on the remaining trajectory pairs after global pruning to complete the similarity incremental update.

[0031] Furthermore, S31 includes the following sub-steps:

[0032] S311. Combine every two trajectories into a trajectory pair and calculate the upper bound of the spatial similarity between the two trajectories in the trajectory pair;

[0033] S312. Calculate the upper bound of the temporal similarity between the two trajectories in the trajectory pair;

[0034] S313. Calculate the upper bound of the spatio-temporal similarity between the two trajectories based on the upper bound of the spatial similarity and the upper bound of the temporal similarity between the two trajectories, and remove the trajectory pair when the upper bound of the spatio-temporal similarity is less than the similarity threshold to complete the global pruning.

[0035] Further, in S311, the trajectory and the trajectory The upper bound of the spatial similarity between The calculation formula is:

[0036] ;

[0037] In the formula, represents the new position, represents the exponent, represents the distance function;

[0038] In S311, the trajectory and the trajectory The upper bound of the temporal similarity between The calculation formula is:

[0039] ;

[0040] In the formula, represents the time function;

[0041] In S313, the trajectory and the trajectory The upper bound of the spatio-temporal similarity between The calculation formula is:

[0042] ;

[0043] In the formula, represents the weight of the trajectory spatial similarity in the trajectory similarity.

[0044] Further, S32 includes the following sub-steps:

[0045] S321. Construct a two-dimensional space plane and divide the two-dimensional space plane into several grid cells;

[0046] S322. With the grid where the new position is located as the center, construct several concentric circles;

[0047] S323. Starting from the smallest concentric circle, conduct a layer-by-layer search, and use the grid where the trajectory position satisfying is located as the candidate grid to complete one search; where, represents the new position, represents the trajectory position, represents the distance function, represents the current search depth, represents the size of the central grid, represents the grid containing the new position;

[0048] S324. In the region other than concentric circles in the two-dimensional space plane, the grids passed by the trajectory are used as candidate grids;

[0049] S325. Among all candidate grids, determine the position closest to the new position to complete the grid neighborhood circular search;

[0050] S326. After completing the grid neighborhood circular search, the trajectory positions that satisfy are removed to complete the similarity incremental update; where represents any point located on the close to the side, represents the position in the trajectory closest to the trajectory before the new position is added, represents the new position, represents and the midpoint between them is located on the perpendicular bisector of the line segment .

[0051] The beneficial effects of the present invention are:

[0052] (1) By introducing a time-aware exponential decay factor into the spatio-temporal similarity function, the present invention ensures real-time dynamic update of the connection results and eliminates outdated results;

[0053] (2) Based on the matrix partitioning scheme and combined with dynamic load balancing, the present invention effectively partitions the trajectory stream while minimizing data redundancy;

[0054] (3) To improve efficiency, the present invention implements a multi-level pruning technique in both spatial and temporal dimensions to accelerate similarity calculation. In addition, the approximate algorithm further optimizes the merging process of the decay similarity. Brief Description of the Drawings

[0055] Figure 1 is a flowchart of the parallel online similarity analysis method for the trajectory stream;

[0056] Figure 2 is an example diagram of the trajectory stream partitioning based on the matrix;

[0057] Figure 3 is an example diagram of the spatial grid neighborhood search;

[0058] Figure 4 is an example diagram of the spatial linear pruning;

[0059] Figure 5 is an influence diagram of the number of trajectory positions |P|;

[0060] Figure 6It is an influence diagram of the number of threads;

[0061] Figure 7 It is an influence diagram of the interval expiration threshold. Detailed implementation manners

[0062] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0063] As Figure 1 shown, the present invention provides a parallel online similarity analysis method for trajectory streams, including the following steps:

[0064] S1. Collect the first trajectory position stream and the second trajectory position stream of moving objects, and perform load partitioning to obtain a plurality of partition matrices;

[0065] S2. Perform load balancing processing on each partition matrix;

[0066] S3. Based on the partition matrices after load balancing processing, perform similarity incremental update;

[0067] S4. Based on the similarity incremental update result, calculate the trajectory similarity using a decay factor when new data is added.

[0068] For load partitioning and dynamic adjustment, a matrix-based partitioning scheme is used to partition the continuously growing trajectory position stream and . Each matrix element corresponds to a parallel thread, which is responsible for downstream similarity update. Assuming the matrix size is , then 9 threads are used to incrementally update the similarity from positions to trajectories.

[0069] To achieve incremental similarity update, a dynamic workload adjustment module is introduced, which allows real-time monitoring and reassigning new data to less-loaded threads when load imbalance occurs. This ensures data independence between partitions and the integrity of the similarity join results. As the input stream changes, the similarity from positions to trajectories is incrementally updated in parallel among partitions. When a new sample position arrives, it triggers multiple updates from positions to trajectories within the partition. To eliminate unqualified trajectory pairs at an early stage, a multi-level pruning technique is implemented, which maintains the upper bound of trajectory spatio-temporal similarity. For trajectory pairs with an upper bound lower than the threshold, global pruning is performed. For the remaining trajectory pairs, a local pruning strategy is adopted, including grid neighborhood search and multi-dimensional linear pruning, to further reduce unnecessary updates.

[0070] Finally, the similarity between the updated position of the upstream and the trajectory is combined into spatio-temporal similarity. Since processing clock updates each time incurs exponential decay and high computational overhead, an approximate algorithm for similarity combination based on interval updates is proposed. It should be noted that the similarity combination between trajectory pairs is independent, so it can be parallelized. Finally, by comparing with , the final connection result can be obtained.

[0071] In the embodiment of the present invention, in S1, the positions of the first trajectory position stream are partitioned in a column-cyclic manner, and the positions of the second trajectory position stream are partitioned in a row-cyclic manner, obtaining a number of partition matrices of size , where represents the number of available threads for downstream trajectory similarity joining.

[0072] Assume is a perfect square number, and a partition matrix of size is constructed, where each element represents a partition. Each partition contains two sets of position streams from and subsets. Allocation is performed with the trajectory as the basic unit to ensure that each partition contains a complete trajectory for further verification. The positions in are allocated in a column-cyclic manner, while the positions in are allocated in a row-cyclic manner. For trajectory , if any position of has been allocated to a partition in a certain column of the matrix, then the subsequent positions of will also be allocated to the partition in that column. Otherwise, the positions of are allocated according to two cases: if the current partition matrix can still maintain load balance, the positions are allocated in a cyclic manner; otherwise, the positions are allocated to the column with the least load.

[0073] Figure 2 Shows an example of partitioning for a matrix size. Let the trajectory set and , where represents trajectory 1, represents trajectory 2, represents trajectory 3, represents trajectory 4, represents trajectory 5, represents trajectory 6, represents trajectory 7, represents trajectory 8, which has been divided into two subsets. Given the trajectory position streams and , where represents the first trajectory point in trajectory 1, represents the first trajectory point in trajectory 2, represents the first trajectory point in trajectory 3, represents the first trajectory point in trajectory 4, represents the first trajectory point in trajectory 5, represents the first trajectory point in trajectory 6, represents the first trajectory point in trajectory 7, represents the first trajectory point in trajectory 8. By connecting and two subsets, a partitioning example can be easily obtained. In the partitioning scheme, each trajectory pair is processed in one and only one partition (for example, Figure 2 in is always verified in the last partition), thus ensuring the independence of calculations.

[0074] The union of the connection results of all partitions is equal to and the complete Cartesian product of the trajectories in, thus ensuring the result integrity. The matrix - based partitioning method proposed by the present invention combines these advantages, thus enabling efficient parallel processing while ensuring the correctness of the results.

[0075] In the embodiment of the present invention, S2 includes the following sub - steps:

[0076] S21. Construct a load matrix with the same size as the partitioning matrix;

[0077] S22. Calculate the first imbalance factor and the second imbalance factor according to the load matrix;

[0078] S23. Perform load balancing processing on the partitioning matrix according to the first imbalance factor and the second imbalance factor.

[0079] In the embodiment of the present invention, although the number of trajectories in each partition is balanced, when there are large differences in the periods of different trajectories, it may lead to uneven load among partitions. However, the periods of trajectories are unpredictable and cannot be determined in advance. To enhance the load balance, a two - stage adjustment algorithm, namely dynamic load monitoring and adjustment, is introduced.

[0080] To continuously monitor the load of each partition, a load matrix with the same size as the partitioning matrix is initialized . Each matrix element records the load of the corresponding partition. When a new position is assigned to a certain partition, the load of that partition increases accordingly. To quantify the partition imbalance of columns and rows, two imbalance factors are defined.

[0081] The imbalance factor changes with the inflow of the input trajectory position. Once the imbalance factor reaches a given threshold , dynamic workload adjustment will be triggered to ensure load balance. Specifically, if the first imbalance factor reaches the threshold , the column with the minimum load will be found . For a new trajectory , the subsequent positions from will be assigned to all partitions in that column. The above assignment operation is repeatedly executed until the first imbalance factor drops below the threshold . If the second imbalance factor reaches , the row with the minimum load will be found to perform the above operation.

[0082] For a new position from , there are three partitioning cases: (i) If has already been assigned to a certain column, then will be assigned to all partitions in that column; otherwise, (ii) If is less than the threshold , then and will be assigned to the columns in a cyclic manner. (iii) If is greater than , then and will be assigned to the column with the minimum load.

[0083] In the embodiment of the present invention, in S22, the calculation formula of the first imbalance factor is:

[0084] ;

[0085] ;

[0086]

[0087] ;

[0088] ; In the formula, represents the total partition load of the -th column in the load matrix, represents the total partition load of the -th row in the load matrix, represents the maximum value operation, represents the minimum value operation, represents the number of available threads for the downstream trajectory similarity join, represents the load at the th row and th column of the load matrix, represents the load at the th row and th column of the load matrix, represents the dimension of the load matrix;

[0089] In S22, the formula for the second imbalance factor is as follows:

[0090] ;

[0091] In the formula, represents the sum of the loads of the th row of the load matrix, represents the sum of the loads of the th row of the load matrix.

[0092] In the embodiment of the present invention, in S23, if the first imbalance factor is greater than or equal to the threshold, the subsequent positions of the first trajectory position stream are assigned to all partition matrices corresponding to the column with the minimum load until the first imbalance factor is less than the threshold;

[0093] In S23, if the second imbalance factor is greater than or equal to the threshold, the subsequent positions of the first trajectory position stream are assigned to all partition matrices corresponding to the row with the minimum load until the second imbalance factor is less than the threshold.

[0094] In the embodiment of the present invention, S3 includes the following sub-steps:

[0095] S31. When adding a new position, perform global pruning on the trajectory pairs;

[0096] S32. Perform local pruning on the remaining trajectory pairs after global pruning to complete the similarity incremental update.

[0097] In the embodiment of the present invention, S31 includes the following sub-steps:

[0098] S311. Combine every two trajectories into a trajectory pair, and calculate the upper bound of the spatial similarity between the two trajectories in the trajectory pair;

[0099] S312. Calculate the upper bound of the temporal similarity between the two trajectories in the trajectory pair;

[0100] S313. According to the upper bound of the spatial similarity and the upper bound of the temporal similarity between the two trajectories, calculate the upper bound of the spatio-temporal similarity between the two trajectories, and when the upper bound of the spatio-temporal similarity is less than the similarity threshold, eliminate the trajectory pair to complete the global pruning.

[0101] In the embodiment of the present invention, in S311, the trajectory The upper bound of the spatial similarity between the and the trajectory is calculated by the formula:

[0102] ;

[0103] In the formula, represents the new position, represents the exponent, represents the distance function;

[0104] In S311, for the upper bound of the temporal similarity between the trajectory and the trajectory is calculated by the formula:

[0105] ;

[0106] In the formula, represents the time function;

[0107] In S313, for the upper bound of the spatio-temporal similarity between the trajectory and the trajectory is calculated by the formula:

[0108] ;

[0109] In the formula,

[0110] represents the weight of the trajectory spatial similarity in the trajectory similarity.

[0111] In the embodiment of the present invention, S32 includes the following sub-steps:

[0112] S321. Construct a two-dimensional spatial plane and divide the two-dimensional spatial plane into a number of grid cells;

[0112] S322. With the grid where the new position is located as the center, construct a number of concentric circles;

[0113] S323. Starting from the smallest concentric circle, perform a layer-by-layer search, and use the grid where the trajectory position satisfying is located as the candidate grid to complete one search; where represents the new position, represents the trajectory position, represents the distance function, represents the current search depth, represents the size of the central grid, represents the grid containing the new position;

[0114] S324. In the region other than concentric circles in the two-dimensional space plane, use the grids passed by the trajectory as candidate grids;

[0115] S325. Among all candidate grids, determine the position closest to the new position to complete the grid neighborhood circular search;

[0116] S326. After completing the grid neighborhood circular search, remove the trajectory positions that satisfy to complete the similarity incremental update; where represents any point located on the close to the side, represents the position in the trajectory closest to the trajectory before the new position is added, represents the new position, represents and the midpoint between is located on the perpendicular bisector of the line segment ;

[0117] In an embodiment of the present invention, once a new position is assigned to a certain column (or row), the trajectory similarity update between and within that column (or row) is triggered. Specifically, two sets of similarity updates are required: (a) the spatial and temporal distances between the position and other trajectories within the partition; (b) the spatial and temporal distances between the sample positions and the trajectory among the trajectories in the trajectory partition except . A direct method is to traverse all sample positions within the partition, but this method is very time-consuming when dealing with a large number of positions. To achieve efficient similarity update, the present invention proposes an incremental similarity update method, which adopts some effective pruning strategies. The present invention also proposes a grid neighborhood circular search and a multi-dimensional linear pruning method, which are used to calculate update groups (a) and (b) respectively.

[0118] In the global pruning strategy, given any two trajectories and , by combining the upper bound of the spatial similarity between the trajectories and the upper bound of the temporal similarity between the trajectories, the upper bound of the spatio-temporal similarity between the two trajectories can be calculated. Using this upper bound, the present invention proposes a global pruning strategy, that is, if the upper bound of the spatio-temporal similarity between two trajectories is less than the threshold, then the trajectory pair It cannot be a connection result, so it can be safely pruned at an early stage .

[0119] Let and represent the same trajectory before and after being recorded at the new position . In the previous similarity calculation, the distances between all positions in the trajectory and all other trajectories within the same partition have been calculated and cached, and can be directly obtained. Assume the number of trajectories in this partition is . The time complexity of obtaining the upper bound is , while calculating the exact similarity requires time. Using the global pruning strategy, trajectories pairs that do not meet the conditions can be safely pruned at an early stage, thus avoiding the exact similarity calculation and improving efficiency. Only for these few remaining candidate trajectories, their exact similarities will be calculated in the subsequent stage.

[0120] To obtain the upper bound , it is necessary to obtain the distances from the position to the trajectories and , where it is necessary to identify the position in the trajectory that is closest to the position and calculate within the same partition.

[0121] In the grid neighborhood circular search, if the new position is recorded, two sets of similarity updates are required. Substantially, the update process of group (a) is the nearest neighbor search from the position to other trajectories within the same partition.

[0122] Given a two-dimensional spatial plane representing the space of trajectory positions, gradually divide this plane into equal-sized grid cells with side length , based on the distribution of the input positions. Initially, consider the first published position as the central position. By expanding from this central position in four directions , a central grid with size is formed. Let represent the boundary of the central grid, where represents the left boundary of the central grid, represents the right boundary of the central grid, represents the lower boundary of the central grid, Represents the upper boundary of the central grid. As new positions are published, it iteratively expands along each axis, starting from the edge of this grid, and finally divides the entire space into grids of equal size. Each position can only belong to one grid.

[0123] Specifically, for the newly generated sample position , the boundaries of the grid it belongs to, denoted as , are calculated as follows:

[0124] ;

[0125] ;

[0126] Given a newly generated position , let be the set of trajectories within the same partition as . To efficiently update , a two-stage method is adopted for calculation. Assume falls within the grid .

[0127] ​​​​​​​​​​​​​​​​​​​​​​​​​, so it can be from In this way, the number of trajectories that need to be further calculated will continue to decrease during the circle search.

[0128] In the second stage, for the trajectories that have not been detected in the previous stage, a one-to-one distance calculation is required to find The minimum distance to these trajectories is usually computationally expensive, especially for long trajectories. To address this problem, a grid-based strategy is used to further prune candidate trajectories. To cover the track A collection of grids with at least one position in . In order to calculate the trajectory from arrive Calculate the distance is the radius of the circular area centered on the circle, ensuring that the circular area covers At least one complete grid in the track Zhongli The closest position must be in a grid that overlaps with the circular area. The present invention records these candidate grids and searches them one by one to find the Recent Locations .

[0129] like Figure 4 As shown, an example of spatial grid neighborhood search is demonstrated with a maximum depth of In the first stage, The first round of search checks the center grid, followed by the circular area formed by the remaining grids. After each round of search, the closest position of the track in these circle grids is found. In the second stage, focus on those tracks that were not detected in the first stage. For long tracks, record the grids that the track passes through as candidate grids (for example, these grids are in the circle and shaded area), where there may be Finally, search among these candidate grids for the nearest The closest location.

[0130] In multidimensional linear pruning, group (b) is updated by effectively reusing the previously computed results. For the new location Before release, track Middle departure track The closest location. and The midpoint between On line segment The perpendicular bisector of This line divides the plane into two regions. Any point near one side , there is , so it can be safely pruned. To achieve this, the present invention first performs coarse-grained pruning at the grid level, removing all grids located away from the nearer side. Then, necessary update verification is performed on the remaining grid positions.

[0131] Taking Figure 4 as an example, it shows how to update when a new position . There are two trajectories and in the figure. It can be observed that before arrives, is the position in trajectory that is closest to , and , while is the position in trajectory that is closest to the remaining positions of trajectory . Here, the vertical lines from to and and are calculated. Only the points located on the right side of and need to be updated (i.e., and ), and the grids in the shaded area can be safely pruned.

[0132] Similarly, for a newly recorded position , the one-dimensional linear pruning method is applied to update the time distance .

[0133] The lower bound and upper bound of the timestamp of that need to be updated are respectively defined by and , where and , represents the other sample positions in trajectory except for the newly recorded position . For the trajectory positions with timestamps less than the lower bound or greater than the upper bound, they can be safely pruned.

[0134] In S4, by introducing exponential decay, whenever an upstream operator transmits a new tuple , the similarity needs to be updated , thus affecting the trajectory and the similarity between. As new data arrives, due to the progress of the system clock, the decay factor for the similarity pairs of all positions in the current partition to the trajectory also changes. Therefore, with the arrival of each new piece of data, it is necessary to traverse the similarity of all positions to the trajectory, recalculate the decay factor, and update the similarity between trajectories. This requires recalculating all trajectory similarities from scratch, which may cause performance bottlenecks.

[0135] Divide the time axis into equal intervals and allocate the data to the corresponding intervals according to the entry time of the data. Each interval uses its middle position as the representative timestamp to replace the independent timestamps of each data point. As time goes by, the decay factor of the earlier intervals approaches zero. To handle data overflow, a very small expiration threshold is set for interval expiration, such as . When the decay factor of an interval drops below this threshold, the interval expires and the data within the interval becomes invalid.

[0136] The decay factor is used to reduce the impact of old data on the current similarity calculation. By making the intervals expire, the impact of early records on trajectory similarity is eliminated. The expiration threshold is adjustable; if set to zero, the expiration mechanism will be disabled, ensuring complete similarity results and comprehensive trajectory evaluation. As each timestamp advances, only the valid intervals need to be updated, thus keeping the number of intervals small and fixed. In contrast, the exact algorithm needs to update , and this value grows with time.

[0137] The present invention will be described below in conjunction with specific embodiments.

[0138] The present invention uses two real-world datasets: the Beijing Taxi Dataset (BTD) and the New York Taxi Dataset (NTD). BTD contains 10,357 trajectories, approximately 17 million sample positions, and the data was collected over 7 days, with 1,686 positions collected per minute. NTD contains 2 million trajectories and approximately 4 million positions, and the data was collected over 45 days, with 61 positions collected per minute. BTD mainly consists of long trajectories, while NTD mainly consists of short trajectories.

[0139] The present invention uses the following metrics to evaluate the efficiency and effectiveness of the model: (a) Average latency, which refers to the average time required to calculate and return the similarity of each pair of trajectories after reading a position. (b) Average throughput, which refers to the number of trajectory pairs processed per unit time. (c) Accuracy, which refers to the spatio-temporal similarity calculated using interval updates compared with the exactly calculated ratio, measuring the accuracy loss caused by interval updates.

[0140] Comparison methods: The present invention evaluates the performance of four methods, specifically: Ori-POTSJ (a direct method for online trajectory similarity join), LP-POTSJ (POTSJ that only uses local pruning techniques), GP-POTSJ (a complete POTSJ framework with global and local pruning techniques), and Ghost* (an extended version of the Ghost framework).

[0141] Efficiency study: That is, the influence of the number of trajectory positions As the number of trajectory positions increases, the cost of calculating the similarity join also rises. To evaluate the efficiency of the model, the present invention changes the size of the dataset and measures the average latency and throughput of the four benchmark methods. Figure 5 (a) and Figure 5 (c) show that in the NTD and BTD datasets, the average latency increases as the dataset size increases, but the increase rate of GP-POTSJ is significantly smaller, which is more than a hundred times more efficient than other benchmark methods. The efficiency of LP-POTSJ is also several times higher than that of Ori-POTSJ and Ghost*, and as increases, the latency grows slowly. Ori-POTSJ and Ghost* perform similarly because they lack optimization in similarity calculation, and their efficiency drops rapidly as the dataset size increases.

[0142] Figure 5 (b) and Figure 5 (d) show that GP-POTSJ is superior to other methods in terms of throughput. However, since short trajectories dominate in NTD, the global and grid-based local pruning methods used by LP-POTSJ and GP-POTSJ have poor effects, resulting in their lower efficiency than when processing BTD.

[0143] Influence of the number of threads: The present invention evaluates the scalability of the framework by changing the number of threads used for computing tasks and measures the average latency and throughput. As Figure 6 shown, increasing the number of threads can reduce latency and increase throughput, indicating that the framework has good scalability.

[0144] Influence of the interval expiration threshold Applying interval updates will result in a loss of accuracy. To evaluate its influence, the present invention fixes the interval length at 1000 milliseconds, varies the interval expiration threshold , and measures its influence on the spatio-temporal similarity accuracy. Ori-POTSJ does not use interval updates, so this experiment focuses on comparing the accuracy of GP-POTSJ and LP-POTSJ. As Figure 7 shown, increasing will result in a slight decrease in accuracy, but the accuracy still remains at Above, this is acceptable.

[0145] Those of ordinary skill in the art will recognize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.

Claims

1. A parallel online similarity analysis method for trajectory streams, characterized in that: The following steps are involved: S1, collecting the first trajectory position flow and the second trajectory position flow of the mobile object, and performing load partitioning to obtain several partition matrices; S2, load balancing processing is performed on each partition matrix; S3, based on the partition matrix after load balancing, perform similarity incremental update; S4. Based on the similarity incremental update results, the trajectory similarity is calculated using the attenuation factor when new data is added.

2. The parallel online similarity analysis method of trajectory flow according to claim 1, characterized in that: In S1, the positions of the first track position stream are partitioned in a column cycle manner, and the positions of the second track position stream are partitioned in a row cycle manner to obtain a plurality of sizes The partition matrix of Indicates the number of threads available for downstream trajectory similarity connections.

3. The parallel online similarity analysis method of trajectory flow according to claim 1, characterized in that: The S2 comprises the following sub-steps: S21, construct a load matrix with the same size as the partition matrix; S22. Calculate a first unbalance factor and a second unbalance factor according to the load matrix; S23: Perform load balancing processing on the partition matrix according to the first imbalance factor and the second imbalance factor.

4. The parallel online similarity analysis method of trajectory flow according to claim 3, characterized in that: In S22, the first unbalance factor The calculation formula is: ; ; ; ; In the formula, Indicates the load matrix The total partition load of the column, Indicates the load matrix The sum of the partition load for the row, Indicates the maximum value operation, represents the minimum operation, Indicates the number of available threads for downstream trajectory similarity connections, Represents the load matrix Line The load of the column, Represents the load matrix Line The load of the column, represents the dimension of the load matrix; In S22, the second unbalance factor The calculation formula is: ; In the formula, Represents the load matrix The total load of the row, Represents the load matrix The total load of the row.

5. The parallel online similarity analysis method of trajectory flow according to claim 3, characterized in that: In S23, if the first imbalance factor is greater than or equal to the threshold, the subsequent positions of the first trajectory position stream are allocated to all partition matrices corresponding to the column with the smallest load until the first imbalance factor is less than the threshold; In S23, if the second imbalance factor is greater than or equal to the threshold, the subsequent positions of the first trajectory position stream are allocated to all partition matrices corresponding to the row with the smallest load until the second imbalance factor is less than the threshold.

6. The parallel online similarity analysis method of trajectory flow according to claim 1, characterized in that: The S3 comprises the following sub-steps: S31, when adding a new position, globally prune the trajectory pair; S32: Perform local pruning on the remaining trajectory pairs after global pruning to complete the similarity incremental update.

7. The parallel online similarity analysis method of trajectory flow according to claim 6, characterized in that: The S31 includes the following sub-steps: S311, forming a trajectory pair with every two trajectories, and calculating the upper bound of the spatial similarity between the two trajectories in the trajectory pair; S312, calculating the upper bound of the time similarity between two trajectories in the trajectory pair; S313: Calculate the upper bound of the spatiotemporal similarity between the two trajectories according to the upper bound of the spatial similarity and the upper bound of the temporal similarity between the two trajectories, and when the upper bound of the spatiotemporal similarity is less than the similarity threshold, remove the trajectory pair to complete global pruning.

8. The parallel online similarity analysis method of trajectory flow according to claim 7, characterized in that: In S311, the trajectory and trajectory The upper bound of the spatial similarity between The calculation formula is: ; In the formula, Indicates the new location, represents the index, represents the distance function; In S311, the trajectory and trajectory The upper bound of the temporal similarity between The calculation formula is: ; In the formula, represents a time function; In S313, the trajectory and trajectory The upper bound of the spatiotemporal similarity between The calculation formula is: ; In the formula, Represents the weight of trajectory space similarity in trajectory similarity.

9. The parallel online similarity analysis method of trajectory flow according to claim 7, characterized in that: The S32 comprises the following sub-steps: S321, constructing a two-dimensional space plane, and dividing the two-dimensional space plane into a plurality of grid units; S322, constructing a number of concentric circles with the grid where the new position is located as the center; S323, starting from the smallest concentric circle, searching layer by layer, The grid where the trajectory position is located is taken as the candidate grid to complete a search; among them, Indicates the new location, represents the trajectory position, represents the distance function, Indicates the current search depth. Indicates the size of the center grid, represents the grid containing the new position; S324, in the area other than the concentric circles in the two-dimensional space plane, taking the grids that the trajectory passes through as candidate grids; S325, determining the position closest to the new position among all candidate grids, and completing the grid neighborhood circular search; S326, after completing the grid neighborhood ring search, The trajectory position of is removed to complete the similarity incremental update; among them, Indicates that it is located near Any point on one side, Indicates adding the previous track to the new position Mid-leave track The nearest location, Indicates the new location, express and The midpoint between The perpendicular bisector of .