Space-time trajectory clustering method based on space-time-POI kernel function
By constructing a space-time-POI kernel function with POI semantic graphs and adaptive dynamic weights, the problems of single feature representation dimensions and noise interference accumulation in trajectory clustering are solved, and trajectory mode mining with high precision and high robustness are achieved, which is suitable for intelligent traffic management and urban planning.
Patent Information
- Application Number
- CN202510823189.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-08
AI Technical Summary
In the existing trajectory clustering methods, the problems of single feature representation dimensions, rigid multi-core fusion mechanisms, and accumulation of noise interference lead to insufficient clustering accuracy and robustness in complex urban market scenarios.
Using a method based on the spatiotemporal-POI kernel function, the POI semantic graph is constructed, node and edge feature vectors are extracted, and adaptive dynamic weights and course learning strategies are combined to integrate spatiotemporal alignment kernels and POI graph structural kernels, optimize clustering centers, suppress noise interference, and improve clustering accuracy and robustness.
It significantly improves the accuracy and robustness of trajectory clustering, is suitable for complex urban market scenarios, provides high-reliability data support, and is suitable for intelligent traffic management and urban planning.
Smart Images

Figure CN120448856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of trajectory clustering technology, and in particular to a spatiotemporal trajectory clustering method based on a spatiotemporal-POI kernel function. Background Art
[0002] With the rapid development of the Internet of Things, mobile communications, and big data technologies, massive amounts of spatiotemporal trajectory data are providing crucial support for intelligent traffic management, travel behavior analysis, and urban planning. Trajectory clustering, as a core technology for mining mobility patterns and identifying hotspots, has a direct impact on the performance of downstream applications due to its accuracy and robustness. Among existing methods for clustering trajectories using kernel functions, most rely on a single kernel function (such as the Gaussian kernel or the cosine kernel) for feature extraction. For example, the Independent Distribution Kernel (IDK) is used to capture the distribution characteristics of trajectories. However, this kernel function only considers the distribution of trajectory points and does not integrate the semantic information of POIs. The data dimension considered is relatively single, and the temporal characteristics of trajectories are ignored. In order to distinguish the direction of trajectories, the temporal attribute is simulated by simply adding the order dimension of trajectory points. This makes it difficult to capture the multi-dimensional characteristics of trajectory data, resulting in a single dimensional feature representation. Furthermore, while existing multi-kernel fusion methods attempt to combine features from multiple kernels, they generally employ a fixed-weight strategy for kernel fusion. This strategy struggles to adapt to the heterogeneity of POI features across different data. Specifically, in areas with high POI density (e.g., commercial centers), the contribution of the POI graph structure kernel cannot be fully amplified, weakening the expressive power of POI features. In areas with low POI density (e.g., remote road sections), excessively high fixed weights introduce irrelevant noise, leading to distorted semantic similarity measurements. At the same time, existing self-clustering methods, such as the self-training-based end-to-end deep trajectory clustering framework E2DTC, consider the impact of high-confidence and low-confidence samples on the clustering effect; the E2DTC framework achieves clustering optimization by aligning the soft assignment results of samples to the target distribution with higher confidence; however, the E2DTC framework involves all samples in the clustering iteration at the initial stage of clustering, which may cause low-confidence samples to interfere with the initial clustering; since low-confidence samples usually contain more noise, the clustering results may be unstable or deviate from the true distribution in the early stage of clustering, thereby exacerbating the accumulation of errors and affecting the convergence speed and clustering accuracy of the model. Summary of the Invention
[0003] In view of this, the present invention proposes a spatiotemporal trajectory clustering method based on the spatiotemporal-POI kernel function to solve the problems of single feature representation dimension, rigid multi-kernel fusion mechanism, and noise interference accumulation in trajectory clustering in complex urban scenes, and to achieve high-precision and high-robustness trajectory pattern mining.
[0004] To achieve the above objectives, the present invention provides a spatiotemporal trajectory clustering method based on a spatiotemporal-POI kernel function, comprising the following steps: S1. Obtain the original trajectory data and perform noise removal and spatiotemporal normalization to obtain the spatiotemporal trajectory sequence. ; S2. Construct a POI semantic graph, and extract node feature vectors and edge feature vectors from the POI semantic graph; S3. Constructing the spatiotemporal alignment kernel and the POI graph structure kernel to extract the spatiotemporal trajectory sequence Features and POI semantic graph features; S4. Calculate the adaptive dynamic weight of the fusion kernel function, adaptively fuse the spatiotemporal alignment kernel and the POI graph structure kernel to form a spatiotemporal-POI function; S401, calculating the POI density of each grid in the POI semantic map; The POI semantic map is divided into M grid areas, and each area is calculated separately. POI density , the expression is: ; S402: Calculate the weighted average of the POI density of all grids where the nodes in the semantic graph POI semantic graph are located as the overall POI density of the POI semantic graph. The expression is: ; in, Represents trajectory POI density of POI semantic graph, trajectory The set of regions divided by the POI semantic graph is , Indicates the maximum POI density value in all areas of the current POI semantic graph. Indicates the minimum POI density value in all areas of the current POI semantic graph; S403: Calculate the average POI density of the trajectory pair , converting the overall POI density into an adaptive dynamic weight through a nonlinear function; The average POI density of the trajectory pair The expression is: ; in, 、 Indicates trajectory; Trajectory pair Adaptive dynamic weight corresponding to POI semantic core The expression is: ; in, Represents the weight coefficient, which linearly scales the average POI density. Indicates the offset; S404: Using adaptive dynamic weights, the spatiotemporal alignment kernel and the POI graph structure kernel are fused to obtain a spatiotemporal-POI kernel function, which is expressed as: ; in, represents the spatiotemporal alignment kernel, Represents the POI graph structure kernel; S5, using the spatiotemporal-POI kernel function to calculate the similarity between trajectory pairs and clustering them using the kernel k-means clustering algorithm; S6. Use the Adam optimizer to optimize and update the cluster centroid, use the gradient descent and cluster loss function to update the parameters of the adaptive dynamic weight, and optimize the adaptive dynamic weight .
[0005] Preferably, the noise removal of the original trajectory data includes removing abnormal drift points, invalid sampling points with empty latitude, longitude or time, and the spatiotemporal standardization process is to convert the timestamp into 24-hour format to generate a standardized spatiotemporal trajectory sequence. , where each trajectory point Includes latitude, longitude and time information.
[0006] Preferably, constructing a POI semantic graph includes the following steps: S201: Divide the geographic coordinate system containing POI information into grids of equal size , grid cell set ; S202. Count each grid The distribution of POI categories within the coverage area is calculated, and the proportion of each category is selected as the grid label. The expression is: ; in, Represents the POI category set within the grid, Indicates the category of POI, Representing a collection All POI categories in Represents the grid Count the POI categories specified in; S203, the space-time trajectory sequence The trajectory points in Map to the grid , the grid center is the node position of the POI semantic graph , the network label is the POI type of the node, and the node set of the POI semantic graph is constructed ; S204: Establish directed connecting edges between adjacent POI nodes according to the trajectory movement sequence to form an edge set, which constitutes the edge set of the POI semantic graph. ; S205: Extracting node feature vectors of the POI semantic graph and edge eigenvectors ; The node feature vector of the POI semantic graph Including node location and POI type, the node location is the longitude of the node ,latitude , change the grid label As the POI type of the node, and using One-hot encoding to represent the POI type of the node, the node feature vector of the POI semantic graph The expression is: ; in, The POI type feature vector representing the node; Edge eigenvector Including the geographical distance between two nodes and normalized angle , use the Haversine formula to calculate the geographical distance between two nodes , the expression is: ; ; ; in, 、 Represents nodes respectively and nodes Latitude, 、 Represents nodes respectively and nodes The difference in longitude and latitude between them; Compute nodes and nodes Azimuth , the expression is: ; The azimuth Convert to angle and normalize to get normalized angle , the expression is: ; The edge eigenvector The expression is: ; Preferably, the spatiotemporal alignment kernel The expression is: ; in, Indicates the calculation of two trajectory sequences The "soft" alignment path between them uses the SoftDTW algorithm to calculate the spatiotemporal similarity between trajectory sequences through a differentiable alignment path; Define the POI graph structure kernel ,The node kernel and edge kernel use the feature vectors of nodes and edges to calculate the similarity of nodes and edges; The POI graph structure core The expression is: ; ; ; ; in, 、 Represents the trajectory 、 POI semantic graph, Representation diagram An edge connecting nodes and nodes , Representation diagram An edge connecting nodes and nodes , Represents a metric node and nodes The node core of similarity between 、 Represents nodes respectively 、 The node feature vector of Represents a metric edge and the edge The edge core of the similarity between 、 Represents edges ,side The edge eigenvector of .
[0007] Preferably, the space-time-POI kernel function is used to calculate the trajectory pair Similarities between , the expression is: ; Clustering using the kernel k-means clustering algorithm includes the following steps: S501, determine the initial cluster centroid, randomly select k center points as the initial cluster centroid ; S502: Introduce a clustering strategy based on curriculum learning and dynamic pseudo-label generation, calculate the confidence of the trajectory and set a confidence threshold, generate dynamic pseudo-labels, and generate high and low confidence labels for the data through the dynamic pseudo-labels and confidence thresholds; S503. Design a clustering loss function that combines course learning and introduce confidence labels into the clustering function; S504: Update the confidence of the data samples that are not involved in the clustering until all the data samples participate in the clustering.
[0008] Preferably, calculating the trajectory confidence comprises the following steps: Calculate the probability of the trajectory being assigned to each cluster and use the t-distribution to measure the probability of each trajectory and cluster centroid Similarities between , the expression is: ; ; in, represents all cluster centroids, Represents the trajectory mapping point to the cluster centroid in the feature space distance, Represents the trajectory mapping point to the cluster centroid in the feature space distance; Calculate the confidence of the trajectory, the expression is: ; ; ; in, Represents trajectory Normalized similarity with N nearest neighbors, representing the trajectory The average similarity of nearby 、 Represents the trajectory The maximum and minimum values of the N nearest neighbor similarity; Setting the confidence threshold , the expression is: ; in, 、 Represent the mean and standard deviation of the confidence of all current samples respectively, represents the control coefficient; According to the confidence threshold Generate trajectory pseudo labels, the expression is: .
[0009] Preferably, the curriculum learning strategy is used to give priority to high-confidence data for trajectory clustering. The clustering loss function is constructed by combining the pseudo-labels of the trajectories with the goals of minimizing the intra-class distance and maximizing the inter-class distance. The expression is: .
[0010] Preferably, the Adam optimizer is used to optimize and update the cluster centroid, and the expression is: ; ; ; ; in, represents the first-order moment estimate of the gradient in momentum form in t iterations, represents the second-order moment estimate of the gradient in momentum form at t iterations, 、 represents the decay coefficient of the two exponentially weighted averages, Represents a constant to avoid division by 0, represents the clustering loss function Cluster centroids Find partial derivatives; After updating the cluster centroid, adjust the threshold value. Second, lower the threshold to include more low-confidence sample data, the expression is: ; in, Indicates the decay rate of the threshold. When the number of iterations hour, Indicates the preset maximum number of iterations, using all samples to cluster and update the centroid; Using gradient descent and clustering loss function Optimize weight parameters w and b to optimize adaptive dynamic weights , the expression is: ; ; in, Represents the learning rate.
[0011] Compared with the prior art, the present invention has the following beneficial effects: The method provided by the present invention maps POI density into dynamic fusion weights through a nonlinear composite function, adaptively adjusting the weight ratio of the spatiotemporal kernel and the semantic kernel according to the POI density distribution. This breaks through the limitations of traditional fixed weight strategies, improves the contribution of the POI graph structure kernel in POI-dense areas, and focuses on the spatiotemporal alignment kernel in sparse areas. This effectively solves the problem of insufficient adaptability of fixed weights in heterogeneous geographical scenarios and significantly enhances the model's generalization ability to complex urban environments. This paper, for the first time, deeply integrates the Soft-DTW-based spatiotemporal alignment kernel with the graph-structured shortest path kernel, simultaneously capturing the temporal characteristics, spatial distribution characteristics, and POI semantic associations of trajectories. This overcomes the shortcomings of existing single-kernel methods, which are single-dimensional and ignore the joint representation of spatiotemporal and POIs, and provides core support for high-precision trajectory clustering. This paper progressively screens samples through a confidence threshold decay mechanism, prioritizing high-confidence data to optimize the centroid and gradually introducing low-confidence samples to suppress noise interference. Compared with the end-to-end framework, the method proposed in this paper can reduce the error accumulation in the early stage of clustering, significantly improve the model convergence speed and robustness, and ensure the stability of the clustering results. It provides high-reliability data support for scenarios such as intelligent traffic management and mobile behavior analysis. The present invention integrates spatiotemporal attributes with POI semantic information, breaks through the limitation of fixed weights, solves the problem of error accumulation caused by interference from low-confidence data in the early stage of clustering, significantly improves clustering accuracy and model robustness, and is suitable for trajectory clustering tasks in complex urban scenarios. It can provide more accurate and reliable data support for downstream applications such as intelligent traffic management, travel behavior analysis, and urban planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a flow chart of a spatiotemporal trajectory clustering method based on a spatiotemporal-POI kernel function according to the present invention; Figure 2 Constructing a POI semantic graph flow chart for the present invention; Figure 3 This is a schematic diagram of the POI semantic graph of the present invention; Figure 4 This is a flow chart of the adaptive dynamic weight calculation of the present invention. DETAILED DESCRIPTION
[0013] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0014] In order to address the problems of single feature representation dimension, rigid multi-core fusion mechanism and accumulated noise interference in trajectory clustering in complex urban scenarios, there is an urgent need for a trajectory clustering method that can realize spatiotemporal-semantic multi-dimensional feature fusion, dynamically adapt to geographical scenarios and effectively suppress noise interference. In recent years, the successful application of curriculum learning strategies in the field of machine learning has shown that the model robustness can be significantly improved through a progressive sample participation mechanism. At the same time, dynamic pseudo-label generation technology provides a new idea for noise sample screening. However, how to organically combine these technologies to construct an adaptive framework suitable for trajectory clustering remains unsolved. Therefore, this embodiment provides a spatiotemporal trajectory clustering method based on the spatiotemporal-POI kernel function, which realizes multi-dimensional feature decoupling modeling through the spatiotemporal-POI kernel function, adapts to POI heterogeneity in combination with an adaptive dynamic weight mechanism, and uses the curriculum learning strategy to progressively optimize the clustering process, thereby breaking through the limitations of existing methods in feature representation, scene adaptability and noise robustness, and providing high-precision clustering support for intelligent traffic management, mobile behavior analysis and other fields. The specific steps include: S1. Obtain the original trajectory data, remove the abnormal drift points, invalid sampling points with empty latitude, longitude or time, and convert the timestamp to 24-hour format to obtain the spatiotemporal trajectory sequence. , where each trajectory point Includes latitude, longitude and time information.
[0015] S2, build POI semantic graph, the flow chart is as follows Figure 2 As shown, a node feature vector and an edge feature vector are extracted from the POI semantic graph; S201: Divide the geographic coordinate system containing POI information into grids of equal size , grid cell set ; S202. Count each grid The distribution of POI categories within the coverage area (such as shopping malls, schools, and residences) is calculated. The type with the highest proportion is selected as the grid label. The expression is: ; in, Represents the POI category set within the grid, Indicates the category of POI, Representing a collection All POI categories in Represents the grid Count the POI categories specified in; S203, the space-time trajectory sequence The trajectory points in Map to the grid , the grid center is the node position of the POI semantic graph , the network label is the POI type of the node, and the node set of the POI semantic graph is constructed ; S204: Establish directed connecting edges between adjacent POI nodes according to the trajectory movement sequence to form an edge set, which constitutes the edge set of the POI semantic graph. , the constructed POI semantic graph is as follows Figure 3 As shown; S205: Extracting node feature vectors of POI semantic graph and edge eigenvectors ; The node feature vector of the POI semantic graph Including node location and POI type, the node location is the longitude of the node ,latitude , change the grid label As the POI type of the node, one-hot encoding is used to represent the POI type of the node. Assuming that there are K types of POIs, the type feature of the node POI is a K-dimensional vector. The node feature vector of the POI semantic graph is The expression is: ; in, Represents the POI type feature vector of the node, , such as the mall type is , school type is ; Edge eigenvector Including the geographical distance between two nodes and normalized angle , use the Haversine formula to calculate the geographical distance between two nodes , the expression is: ; ; ; in, 、 Represents nodes respectively and nodes Latitude, 、 Represents nodes respectively and nodes The difference in longitude and latitude between them; Compute nodes and nodes Azimuth , the expression is: ; The azimuth Convert to angle and normalize to get normalized angle , the expression is: ; Edge eigenvector The expression is: .
[0016] S3, construct the spatiotemporal alignment kernel and the POI graph structure kernel, the spatiotemporal alignment kernel The expression is: ; in, Indicates the calculation of two trajectory sequences The "soft" alignment path between them adopts the SoftDTW algorithm, introduces a smooth minimization operation to make the alignment process differentiable, and calculates the spatiotemporal similarity between trajectory sequences through the differentiable alignment path. The core idea of is to use dynamic programming to calculate the weighted sum of the costs of all possible alignment paths, rather than just selecting the minimum cost path in traditional DTW. Therefore, in this embodiment, The expression is: ; in, Represents the cumulative cost matrix, used to calculate the trajectory and trajectory The cost of each local point pair during the alignment process is calculated, and the final alignment cost is recursively solved through dynamic programming. The expression is: ; ; in, represents the local distance in space and time between two points in the trajectory sequence, Represents the calculation method of the soft minimum in SoftDTW, which can better capture time deformation (such as local stretching, translation, etc.), making the algorithm smoother and more robust when dealing with noise and local disturbances. Indicates the dynamic adjustment of the stringency of trajectory sequence alignment. When it approaches 0, SoftDTW will degenerate into the classic DTW algorithm, which is expressed as: ; Define the POI graph structure kernel ,The node kernel and edge kernel use the feature vectors of nodes and edges to calculate the similarity of nodes and edges; POI graph structure kernel The expression is: ; ; ; ; in, 、 Represents the trajectory 、 POI semantic graph, Representation diagram An edge connecting nodes and nodes , Representation diagram An edge connecting nodes and nodes , Represents a metric node and nodes The node core of similarity between 、 Represents nodes respectively 、 The node feature vector of Represents a metric edge and the edge The edge core of the similarity between 、 Represents edges ,side The edge feature vector of The similarity of graphs can be measured by combining the starting point, end point, and edges; According to the spatiotemporal alignment kernel and POI graph structure kernel Extracting spatiotemporal trajectory sequences Features and POI semantic graph features; Since most existing methods for clustering trajectories using kernel functions use single-kernel architectures such as Gaussian kernels and cosine kernels to extract trajectory features, for example, a trajectory clustering method based on independent distribution kernels (TIDKC) uses the independent distribution kernel (IDK) to capture the distribution features of trajectories. However, this kernel function only considers the distribution of trajectory points and does not integrate the semantic information of POIs. The data dimension considered is relatively single. At the same time, TIDKC does not consider the temporal information of the trajectory. In order to distinguish the direction of the trajectory, it simply adds a dimension to the data dimension to represent the order of each trajectory point. To address the above problems, this embodiment designs a spatiotemporal alignment kernel and a POI graph structure kernel in step S3. For the first time, the kernel function designed based on the SoftDTW algorithm is combined with the graph-based shortest path kernel to capture spatiotemporal trajectory sequence features and POI semantic features. At the same time, the temporal attributes, spatial attributes, and POI semantic information of the trajectory are considered from the perspective of extracting trajectory features and POI semantic features from multiple data dimensions and multiple kernels. This solves the problems of existing kernel function trajectory clustering methods that use a single kernel to capture trajectory features, do not integrate POI semantic information, and do not consider the temporal attributes of trajectory points.
[0017] S4. Calculate the adaptive dynamic weight of the fusion kernel function, adaptively fuse the spatiotemporal alignment kernel and the POI graph structure kernel to form a spatiotemporal-POI function; S401, calculating the POI density of each grid in the POI semantic map; Divide the POI semantic map into M grid areas and calculate each area separately POI density , the expression is: ; S402: Calculate the weighted average of the POI density of all the grids where the nodes are located in the semantic graph POI semantic graph as the overall POI density of the POI semantic graph. The expression is: ; in, Represents trajectory POI density of POI semantic graph, trajectory The set of regions divided by the POI semantic graph is , Indicates the maximum POI density value in all areas of the current POI semantic graph. Indicates the minimum POI density value in all areas of the current POI semantic graph; S403: Calculate the average POI density of the trajectory pair , converting the overall POI density into an adaptive dynamic weight through a nonlinear function; Average POI density of trajectory pairs The expression is: ; in, 、 Indicates trajectory; Trajectory pair Adaptive dynamic weight corresponding to POI semantic core The expression is: ; in, Represents the weight coefficient, which linearly scales the average POI density. Indicates the offset; S404: Using adaptive dynamic weights, the spatiotemporal alignment kernel and the POI graph structure kernel are fused to obtain a spatiotemporal-POI kernel function, which is expressed as: ; in, represents the spatiotemporal alignment kernel, Represents the POI graph structure kernel; Existing multi-kernel fusion methods generally use a fixed weight strategy to fuse kernel functions. This fixed weight strategy is difficult to adapt to the heterogeneity of POI features in different data. In areas with high POI density (such as commercial centers), the fixed weight strategy cannot fully amplify the feature contribution of the POI graph structure kernel, resulting in distortion of the semantic similarity measurement. In areas with low POI density (such as remote road sections), excessively high fixed weights will introduce irrelevant noise. To address the above problems, this embodiment proposes an adaptive dynamic weight in step S4. The core of this method is to automatically adjust the weight of the kernel function fusion according to the POI density. For example, when the POI density is high, the fusion weight of the POI graph structure kernel in the spatiotemporal-POI kernel function is increased, and vice versa. The adaptive dynamic weight balances the fusion ratio of the spatiotemporal alignment kernel and the POI graph structure kernel, thereby effectively maintaining the integrity of the trajectory sequence features and POI semantic feature representations in different scenarios, improving scenario adaptability, avoiding the use of fixed weights that may lead to the loss of features captured by the kernel function, and improving the generalization ability of the model.
[0018] S5, using the spatiotemporal-POI kernel function to calculate the similarity between trajectory pairs and clustering them using the kernel k-means clustering algorithm; Calculate trajectory pairs using the spatiotemporal-POI kernel function Similarities between , the expression is: ; Clustering using the kernel k-means clustering algorithm includes the following steps: S501, determine the initial cluster centroid, randomly select k center points as the initial cluster centroid ; S502: Introduce a clustering strategy based on curriculum learning and dynamic pseudo-label generation, calculate the confidence of the trajectory and set a confidence threshold, generate dynamic pseudo-labels, and generate high and low confidence labels for the data through the dynamic pseudo-labels and confidence thresholds; Calculating trajectory confidence involves the following steps: Calculate the probability of the trajectory being assigned to each cluster and use the t-distribution to measure the probability of each trajectory and cluster centroid Similarities between , the expression is: ; ; in, represents all cluster centroids, Represents the trajectory mapping point to the cluster centroid in the feature space distance, Represents the trajectory mapping point to the cluster centroid in the feature space distance; Calculate the confidence of the trajectory, the expression is: ; ; ; in, Represents trajectory Normalized similarity with N nearest neighbors, representing the trajectory The average similarity of nearby 、 Represents the trajectory The maximum and minimum values of the N nearest neighbor similarity; Setting the confidence threshold , the expression is: ; in, 、 Represent the mean and standard deviation of the confidence of all current samples respectively, represents the control coefficient; According to the confidence threshold Generate trajectory pseudo labels, the expression is: ; S503. Design a clustering loss function based on course learning, give priority to high-confidence data for trajectory clustering, construct a clustering loss function by combining the pseudo-labels of the trajectories with the goals of minimizing the intra-class distance and maximizing the inter-class distance, and introduce confidence labels into the clustering function to ensure that high-confidence samples are preferentially clustered; clustering loss function The expression is: ; S504: Update the confidence of the data samples that are not included in the clustering. As the clustering iteration proceeds, gradually lower the confidence threshold until all data samples participate in the clustering. In the existing self-clustering stage, for example, the self-training-based end-to-end deep trajectory clustering framework E2DTC has considered the impact of high-confidence and low-confidence samples on the clustering effect. E2DTC achieves clustering optimization by aligning the soft assignment results of the samples to the target distribution with higher confidence. However, this method introduces all samples to participate in the clustering iteration at the initial stage of clustering, which may cause low-confidence samples to interfere with the initial clustering. Since low-confidence samples usually contain more noise, they may cause instability of the clustering results or deviate from the true distribution in the early stage of clustering. This aggravates the accumulation of errors and affects the convergence speed and clustering accuracy of the model. To address the above problems, this embodiment proposes a clustering strategy that integrates course learning and dynamic pseudo-label generation in step S5, and introduces dynamic pseudo-labels into the clustering loss function. This strategy sets a confidence threshold to prioritize high-confidence samples in clustering training. As clustering iterations proceed, the confidence threshold is gradually lowered, and low-confidence samples are introduced in an orderly manner. This can effectively reduce the interference of noise samples on model training and reduce error accumulation, thereby significantly improving the robustness of the model and the stability of the clustering results.
[0019] S6. Use the Adam optimizer to optimize and update the cluster centroid. The expression is: ; ; ; ; in, represents the first-order moment estimate of the gradient in momentum form in t iterations, represents the second-order moment estimate of the gradient in momentum form at t iterations, 、 represents the decay coefficient of the two exponentially weighted averages, Represents a constant to avoid division by 0, represents the clustering loss function Cluster centroids Find partial derivatives; After updating the cluster centroid, adjust the threshold value. Second, lower the threshold to include more low-confidence sample data, the expression is: ; in, Indicates the decay rate of the threshold. When the number of iterations hour, Indicates the preset maximum number of iterations, using all samples to cluster and update the centroid; Using gradient descent and clustering loss function Optimize weight parameters w and b, optimize adaptive dynamic weights , the expression is: ; ; in, Represents the learning rate.
[0020] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A spatiotemporal trajectory clustering method based on spatiotemporal-POI kernel function, characterized by: The following steps are involved: S1. Obtain the original trajectory data and perform noise removal and spatiotemporal normalization to obtain the spatiotemporal trajectory sequence. ; S2. Construct a POI semantic graph, and extract node feature vectors and edge feature vectors from the POI semantic graph; S3. Constructing the spatiotemporal alignment kernel and the POI graph structure kernel to extract the spatiotemporal trajectory sequence Features and POI semantic graph features; S4. Calculate the adaptive dynamic weight of the fusion kernel function, adaptively fuse the spatiotemporal alignment kernel and the POI graph structure kernel to form a spatiotemporal-POI function; S401, calculating the POI density of each grid in the POI semantic map; The POI semantic map is divided into M grid areas, and each area is calculated separately. POI density , the expression is: ; S402: Calculate the weighted average of the POI density of all grids where the nodes in the semantic graph POI semantic graph are located as the overall POI density of the POI semantic graph. The expression is: ; in, Represents trajectory POI density of POI semantic graph, trajectory The set of regions divided by the POI semantic graph is , Indicates the maximum POI density value in all areas of the current POI semantic graph. Indicates the minimum POI density value in all areas of the current POI semantic graph; S403: Calculate the average POI density of the trajectory pair , converting the overall POI density into an adaptive dynamic weight through a nonlinear function; The average POI density of the trajectory pair The expression is: ; in, 、 Indicates trajectory; Trajectory pair Adaptive dynamic weight corresponding to POI semantic core The expression is: ; in, Represents the weight coefficient, which linearly scales the average POI density. Indicates the offset; S404: Using adaptive dynamic weights, the spatiotemporal alignment kernel and the POI graph structure kernel are fused to obtain a spatiotemporal-POI kernel function, which is expressed as: ; in, represents the spatiotemporal alignment kernel, Represents the POI graph structure kernel; S5, using the spatiotemporal-POI kernel function to calculate the similarity between trajectory pairs and clustering them using the kernel k-means clustering algorithm; S6. Use the Adam optimizer to optimize and update the cluster centroid, use the gradient descent and cluster loss function to update the parameters of the adaptive dynamic weight, and optimize the adaptive dynamic weight .
2. The spatiotemporal trajectory clustering method based on the spatiotemporal-POI kernel function according to claim 1, characterized in that: The noise removal of the original trajectory data includes removing abnormal drift points, invalid sampling points with empty latitude, longitude or time. The spatiotemporal standardization process converts the timestamp into 24-hour format to generate a standardized spatiotemporal trajectory sequence. , where each trajectory point Includes latitude, longitude and time information.
3. The spatiotemporal trajectory clustering method based on the spatiotemporal-POI kernel function according to claim 1, characterized in that: Constructing a POI semantic graph includes the following steps: S201: Divide the geographic coordinate system containing POI information into grids of equal size , grid cell set ; S202. Count each grid The distribution of POI categories within the coverage area is calculated, and the proportion of each category is selected as the grid label. The expression is: ; in, Represents the POI category set within the grid, Indicates the category of POI, Representing a collection All POI categories in Represents the grid Count the POI categories specified in; S203, the space-time trajectory sequence The trajectory points in Map to the grid , the grid center is the node position of the POI semantic graph , the network label is the POI type of the node, and the node set of the POI semantic graph is constructed ; S204: Establish directed connecting edges between adjacent POI nodes according to the trajectory movement sequence to form an edge set, which constitutes the edge set of the POI semantic graph. ; S205: Extracting node feature vectors of the POI semantic graph and edge eigenvectors ; The node feature vector of the POI semantic graph Including node location and POI type, the node location is the longitude of the node ,latitude , change the grid label As the POI type of the node, and using One-hot encoding to represent the POI type of the node, the node feature vector of the POI semantic graph The expression is: ; in, The POI type feature vector representing the node; Edge eigenvector Including the geographical distance between two nodes and normalized angle , use the Haversine formula to calculate the geographical distance between two nodes , the expression is: ; ; ; in, 、 Represents nodes respectively and nodes Latitude, 、 Represents nodes respectively and nodes The difference in longitude and latitude between them; Compute nodes and nodes Azimuth , the expression is: ; The azimuth Convert to angle and normalize to get normalized angle , the expression is: ; The edge eigenvector The expression is: 。 4. The spatiotemporal trajectory clustering method based on spatiotemporal-POI kernel function according to claim 1, characterized in that: Spatiotemporal alignment kernel The expression is: ; in, Indicates the calculation of two trajectory sequences The "soft" alignment path between them uses the SoftDTW algorithm to calculate the spatiotemporal similarity between trajectory sequences through a differentiable alignment path; Define the POI graph structure core ,The node kernel and edge kernel use the feature vectors of nodes and edges to calculate the similarity of nodes and edges; The POI graph structure core The expression is: ; ; ; ; in, 、 Represents the trajectory 、 POI semantic graph, Representation diagram An edge connecting nodes and nodes , Representation diagram An edge connecting nodes and nodes , Represents a metric node and nodes The node core of similarity between 、 Represents nodes respectively 、 The node feature vector of Represents a metric edge and the edge The edge core of the similarity between 、 Represents edges ,side The edge eigenvector of .
5. The spatiotemporal trajectory clustering method based on spatiotemporal-POI kernel function according to claim 1, characterized in that: Calculate trajectory pairs using the spatiotemporal-POI kernel function Similarities between , the expression is: ; Clustering using the kernel k-means clustering algorithm includes the following steps: S501, determine the initial cluster centroid, randomly select k center points as the initial cluster centroid ; S502: Introduce a clustering strategy based on curriculum learning and dynamic pseudo-label generation, calculate the confidence of the trajectory and set a confidence threshold, generate dynamic pseudo-labels, and generate high and low confidence labels for the data through the dynamic pseudo-labels and confidence thresholds; S503. Design a clustering loss function that combines course learning and introduce confidence labels into the clustering function; S504: Update the confidence of the data samples that are not involved in the clustering until all the data samples participate in the clustering.
6. The spatiotemporal trajectory clustering method based on spatiotemporal-POI kernel function according to claim 5, characterized in that: Calculating trajectory confidence involves the following steps: Calculate the probability of the trajectory being assigned to each cluster and use the t-distribution to measure the probability of each trajectory and cluster centroid Similarities between , the expression is: ; ; in, represents all cluster centroids, Represents the trajectory mapping point to the cluster centroid in the feature space distance, Represents the trajectory mapping point to the cluster centroid in the feature space distance; Calculate the confidence of the trajectory, the expression is: ; ; ; in, Represents trajectory Normalized similarity with N nearest neighbors, representing the trajectory The average similarity of nearby 、 Represents the trajectory The maximum and minimum values of the N nearest neighbor similarity; Setting the confidence threshold , the expression is: ; in, 、 Represent the mean and standard deviation of the confidence of all current samples respectively, represents the control coefficient; According to the confidence threshold Generate trajectory pseudo labels, the expression is: 。 7. The spatiotemporal trajectory clustering method based on spatiotemporal-POI kernel function according to claim 6, characterized in that: Using the curriculum learning strategy, we prioritize high-confidence data for trajectory clustering. We construct a clustering loss function by combining the pseudo-labels of the trajectories with the goal of minimizing the intra-class distance and maximizing the inter-class distance. The expression is: 。 8. The spatiotemporal trajectory clustering method based on spatiotemporal-POI kernel function according to claim 1, characterized in that: The Adam optimizer is used to optimize and update the cluster centroid. The expression is: ; ; ; ; in, represents the first-order moment estimate of the gradient in momentum form in t iterations, represents the second-order moment estimate of the gradient in momentum form at t iterations, 、 represents the decay coefficient of the two exponentially weighted averages, Represents a constant to avoid division by 0, represents the clustering loss function Cluster centroids Find partial derivatives; After updating the cluster centroid, adjust the threshold value. Second, lower the threshold to include more low-confidence sample data, the expression is: ; in, Indicates the decay rate of the threshold. When the number of iterations hour, Indicates the preset maximum number of iterations, using all samples to cluster and update the centroid; Using gradient descent and clustering loss function Optimize weight parameters w and b to optimize adaptive dynamic weights , the expression is: ; ; in, Represents the learning rate.